Large model layered caching loading method oriented to Plevatory capacity intelligent computing center

By disassembling the big model into multiple models and hierarchically and cached it to local storage devices according to the number of usages, the problems of low loading efficiency and high cost of large models in the intelligent computing center are solved, and efficient model loading and the application of universal computing power are achieved.

CN120216463AActive Publication Date: 2025-06-27DATACANVAS LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510694269.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-06-27
Estimated Expiration
2045-05-27

AI Technical Summary

Technical Problem

Intelligent computing centers need to load large models from remote databases, resulting in high economic costs and low loading efficiency, making it difficult to achieve widespread application of universal computing power.

Method used

By disassembling the large model into multiple model hierarchies, and counting the number of usages of these hierarchies in real time, cache them to the storage device of the local computing node according to the number of usages, so that they can be loaded directly from the local area when they need to be loaded.

Benefits of technology

It effectively saves network traffic used for large-scale model loading, improves loading speed, reduces economic costs, and promotes the widespread application of universal computing power.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216463A_ABST
    Figure CN120216463A_ABST
Patent Text Reader

Abstract

The invention provides a large model hierarchical cache loading method oriented to a Plevatory computing power intelligent computing center, and relates to the technical field of intelligent computing centers, intelligent computing centers and computing power infrastructures, the method comprises the following steps: S1, counting the number of use times of model hierarchies stored in a remote database, the model hierarchies being obtained by disassembling a large model; s2, when the number of use times of the model layer exceeds a first preset threshold value, copying the model layer from a remote database to a local first storage device; s3, when the number of use times of the model hierarchy exceeds a second preset threshold value, whether the model hierarchy is a public model hierarchy or a private model hierarchy is judged, if the model hierarchy is the public model hierarchy, the model hierarchy is moved from the first storage device to a local second storage device, and the second preset threshold value is larger than the first preset threshold value; the access speed of the second storage device is higher than that of the first storage device. According to the invention, the network flow used for large model loading can be saved and the loading speed can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of intelligent computing centers, intelligent computing centers, and computing power infrastructure, and specifically relates to a large model hierarchical caching and loading method for an intelligent computing center for inclusive computing power. Background Art

[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "intelligent computing centers" have emerged as the times require.

[0003] An "intelligent computing center" refers to a facility that provides the required computing power, data, and algorithms for artificial intelligence applications (such as scenarios like artificial intelligence deep learning model development, model training, and model inference) by using large-scale heterogeneous computing power resources, including general computing power and intelligent computing power. The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.

[0004] The "intelligent computing center" includes but is not limited to the "intelligent computing center".

[0005] An "intelligent computing center", that is, an artificial intelligence computing center, is a type of computing power infrastructure that provides computing power services, data services, and algorithm services required for artificial intelligence applications based on artificial intelligence theory and using an artificial intelligence computing architecture.

[0006] "Computing power" is the core of "intelligent computing centers" and "intelligent computing centers". It is the ability of computer devices or computing / data centers to process information, the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, the computing ability to achieve the output of target results by processing information data, and a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity. It mainly provides services to society through computing power infrastructure.

[0007] An intelligent computing center can provide data storage capabilities for storing large models of users of tenant services. Currently, large models are usually stored in a remote database. When a large model needs to be read, it needs to be loaded from the remote database to the local. Large models can be as small as several G or as large as hundreds of G. Each time loading from the remote database will bring a large network traffic overhead. In addition, the loading speed will also be affected by factors such as network speed and the storage performance of computing power resources. That is, the existing large model loading methods have high economic costs and low loading efficiency, and it is difficult to achieve the wide application of inclusive computing power. Summary of the Invention

[0008] The present invention provides a large model hierarchical caching and loading method for an intelligent computing center for inclusive computing power, which is used to solve the problems that an intelligent computing center needs to load a large model from a remote database, has high economic costs, low loading efficiency, and it is difficult to achieve the wide application of inclusive computing power.

[0009] To solve the above technical problems, the present invention is implemented as follows: In a first aspect, the present invention provides a large model hierarchical caching and loading method for an inclusive computing power intelligent computing center, including: Step S1: Real-time statistics of the usage times of the model hierarchies stored in the remote database, where the model hierarchies are obtained by disassembling the large model, and each large model stored in the remote database is disassembled into multiple model hierarchies for storage; Step S2: When the usage times of the model hierarchy exceed a first preset threshold, copy the model hierarchy from the remote database to the first storage device of the local computing power node; Step S3: When the usage times of the model hierarchy exceed a second preset threshold, determine whether the model hierarchy is a public model hierarchy or a private model hierarchy. If the model hierarchy is a public model hierarchy, move the model hierarchy from the first storage device to the second storage device of the local computing power node, where the second preset threshold is greater than the first preset threshold, the access speed of the second storage device is higher than that of the first storage device, the public model hierarchy refers to the model hierarchy shared by multiple tenants, and the private model hierarchy refers to the model hierarchy that only belongs to one tenant.

[0010] Optionally, step S2 includes: Step S21: When the usage times of the model hierarchy exceed a first preset threshold, copy the model hierarchy from the remote database to the first storage devices of all local computing power nodes of the intelligent computing center; among them, for the local computing power node where the user currently needs to read the model hierarchy, copy the model hierarchy from the remote database to the first storage device of the local computing power node; for other local computing power nodes of the intelligent computing center, monitor the usage ratio of the network bandwidth between the other local computing power nodes and the remote database, and when the usage ratio of the network bandwidth is less than the preset bandwidth usage ratio threshold, copy the model hierarchy from the remote database to the first storage device of the other local computing power nodes; And / or Step S3 includes: Step S31: When the usage times of the model hierarchy exceed a second preset threshold, determine whether the model hierarchy is a public model hierarchy or a private model hierarchy. If the model hierarchy is a public model hierarchy, for each local computing power node of the intelligent computing center, move the model hierarchy from the first storage device of the local computing power node to the second storage device of the local computing power node.

[0011] Optionally, if the large model is a model obtained by fine-tuning a base model, the large model includes two model layers, and the two model layers include: a base model parameter layer and a fine-tuning parameter layer; and / or The large model includes multiple model layers disassembled based on the model structure.

[0012] Optionally, the method further includes: Step S4: Monitor the storage capacity of the second storage device in real time. If the storage capacity of the second storage device is higher than the first preset capacity ratio, compare the usage times of each of the common model layers stored in the second storage device, and move the common model layer with the lowest usage times to the first storage device for storage.

[0013] Optionally, the method further includes: Step S5: Monitor the storage capacity of the first storage device in real time. If the storage capacity of the first storage device is higher than the second preset capacity ratio, determine whether the private model layer is stored in the first storage device. If the private model layer is stored in the first storage device, compare the usage times of each of the private model layers stored in the first storage device, and delete the private model layer with the lowest usage times; if the private model layer is not stored in the first storage device, compare the usage times of each of the common model layers stored in the first storage device, and delete the common model layer with the lowest usage times.

[0014] Optionally, the method further includes: Step S6: In response to an instruction triggered by the user to read the large model, obtain information about all model layers of the large model; Step S7: Determine whether the model layer is a common model layer or a private model layer according to the information of the model layer; Step S8: If the model layer is a common model layer, query whether the model layer is stored in the second storage device. If the model layer is stored in the second storage device, load the model layer into the memory; if the model layer is not stored in the second storage device, query whether the model layer is stored in the first storage device; if the model layer is stored in the first storage device, load the model layer into the memory, and if the model layer is not stored in the first storage device, pull the model layer from the remote database and save it to the first storage device, and load the model layer into the memory; Step S9: If the model layer is a private model layer, query whether the model layer is stored in the first storage device; if the model layer is stored in the first storage device, load the model layer into memory, and if the model layer is not stored in the first storage device, pull the model layer from the remote database, save it to the first storage device, and load the model layer into memory; Step S10: Merge all the model layers of the large model loaded into the memory to obtain the large model.

[0015] Optionally, the method further includes: Step S11: Each time the model layer is used by a user of the tenant, increment the usage count of the model layer by 1.

[0016] Optionally, before step S1, it further includes: Step S12: In response to an instruction to save the large model triggered by the user, disassemble the user's large model to obtain multiple model layers; Step S13: Calculate the hash value of the content of each model layer; Step S14: Retrieve whether there is an index record corresponding to the hash value of the model layer in the index library. If it exists, do not save the model layer; if it does not exist, save the model layer to the remote database and generate an index record based on the hash value of the model layer and the storage path of the model layer and store it in the index library; Step S15: Record the index values of all model layers of the large model.

[0017] Optionally, step S14 includes: Step S141: Use the hash value of the model layer as the index value, generate an index record based on the index value of the model layer and the storage path of the model layer, and store the generated index record in the index library.

[0018] Optionally, step S14 further includes: Step S142: If there is no index record corresponding to the hash value of the model layer in the index library, save the model layer as a private model layer to the private database of the remote database, generate an index record corresponding to the hash value of the model layer and store it in the index library, and record the tenant to which the user belongs in the index record; and / or Step S143: If there is an index record corresponding to the hash value of the model layer in the index library, determine whether the tenant to which the user belongs is recorded in the index record; Step S144: If the tenant to which the user belongs is not recorded in the index record, increment the tenant sharing count in the index record by 1 and record the tenant to which the user belongs in the index record; determine whether the model layer is stored in the private database. If the model layer is stored in the private database, move the model layer as a public model layer to the public database of the remote database. Step S145: If the tenant to which the user belongs is already recorded in the index record, there is no need to increment the tenant sharing count in the index record.

[0019] In a second aspect, the present invention provides a large model layer caching and loading device for an inclusive computing power intelligent computing center, including: A statistics module for real-time statistics of the usage times of the model layers stored in the remote database, where the model layers are obtained by disassembling a large model, and each large model stored in the remote database is disassembled into multiple model layers for storage. A first caching module for copying the model layer from the remote database to the first storage device of the local computing power node when the usage times of the model layer exceed a first preset threshold. A second caching module for determining whether the model layer is a public model layer or a private model layer when the usage times of the model layer exceed a second preset threshold. If the model layer is a public model layer, move the model layer from the first storage device to the second storage device of the local computing power node, where the second preset threshold is greater than the first preset threshold, the access speed of the second storage device is higher than that of the first storage device, the public model layer refers to the model layer shared by multiple tenants, and the private model layer refers to the model layer that only belongs to one tenant.

[0020] In a third aspect, the present invention provides a computing power device, including: a processor, a memory, and a program stored on the memory and executable on the processor. When the program is executed by the processor, the steps of the large model layer caching and loading method for an inclusive computing power intelligent computing center as described in the first aspect above are implemented.

[0021] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the large model layer caching and loading method for an inclusive computing power intelligent computing center as described in the first aspect above are implemented.

[0022] Fifth aspect, the present invention provides a computer program product, including computer instructions, which when executed by a processor implement the steps of the large model hierarchical caching and loading method for the inclusive computing power intelligent computing center as described in the first aspect above.

[0023] In the present invention, the large model of the user is disassembled to obtain multiple model layers, which are stored in a remote database. The usage times of the model layers stored in the remote database are counted in real time. According to the usage times of the model layers, the model layers are cached to the storage device of the local computing power node. Thus, when it is necessary to load a model layer, if the model layer is stored in the storage device of the local computing power node, it can be directly loaded from the storage device of the local computing power node into the memory, which can effectively save the network traffic used for loading the large model, and effectively improve the loading speed of the large model, greatly reduce the economic cost, and promote the wide application of inclusive computing power. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings: Figure 1 is a schematic flow chart of the large model hierarchical caching and loading method for the inclusive computing power intelligent computing center of the present invention; Figure 2 is a schematic structural diagram of the large model hierarchical caching and loading device for the inclusive computing power intelligent computing center of the present invention; Figure 3 is a schematic structural diagram of the computing power device of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] The technical solutions in the present invention will be clearly and completely described below with reference to the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.

[0026] First, the technical terms related to the present invention will be briefly described below.

[0027] The "computing power" as described in the present invention refers to: the ability of a computer device or a computing / data center to process information, which is the ability of computer hardware and software to cooperate to execute a certain computing requirement, the computing ability to process information data and output a target result, and a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, and mainly provides services to society through computing power infrastructure.

[0028] The "computational power" (CP) as described in the present invention refers to: the ability of a data center server to process data and output results, which is a comprehensive indicator for measuring the computing ability of a data center and includes general computing ability, supercomputing ability, and intelligent computing ability. The commonly used measurement unit is the number of floating-point operations per second (FLOPS, 1 EFLOPS = 10^18 FLOPS), and the larger the value, the stronger the comprehensive computing ability. It is estimated that 1 EFLOPS is approximately the computing power output of 5 Tianhe-2A or 500,000 mainstream server CPUs or 2 million mainstream laptops. The calculation formula is: CP = CP 通用 + CP 智能 + CP 超级 。

[0029] The "carrying capacity" (NP) as described in the present invention refers to: the performance of the data transmission ability of computing power facilities, which is a comprehensive ability including network architecture, network bandwidth, transmission delay, intelligent management and scheduling, etc., and involves network transmission inside and between data centers, and is a comprehensive indicator for measuring network transmission scheduling ability.

[0030] The "storage power" (SP) as described in the present invention refers to: the comprehensive ability of a data center in four aspects: data storage capacity, performance, security and reliability, and green and low-carbon, which is a comprehensive indicator for measuring the data storage ability of a data center and includes external storage devices such as storage arrays and server internal storage devices. The commonly used measurement unit for storage capacity is exabyte (EB, 1 EB = 2^60 bytes), the commonly used measurement unit for performance is the number of read / write operations per second per unit capacity (IOPS / TB, Input / Output Operations Per Second / TB), and the disaster recovery ratio is an important manifestation of security and reliability.

[0031] The "computing power infrastructure" as described in the present invention refers to: a new type of information infrastructure integrating information computing power, network carrying capacity, and data storage capacity, which can realize the centralized computing, storage, transmission, and application of information.

[0032] The "new information infrastructure" described in the present invention refers to: mainly including network infrastructures such as 5G networks, fiber broadband networks, backbone networks, international communication networks, satellite Internet, etc., computing power infrastructures such as data centers, general computing power centers, intelligent computing centers, supercomputing centers, etc., and new technology facilities such as artificial intelligence, blockchain, and quantum computing.

[0033] The "computing power" described in the present invention includes: general computing power, intelligent computing power, and super computing power.

[0034] The "general computing power" described in the present invention refers to: the computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.

[0035] The "intelligent computing power" described in the present invention refers to: for various artificial intelligence innovation applications, a computing platform is deployed on a large scale based on dedicated chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), and ASIC (Application Specific Integrated Circuit), such as natural language processing, machine vision, and so on.

[0036] The "super computing power" described in the present invention refers to: mainly the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and processes extremely complex or data-intensive problems through a dedicated operating system. It is mainly used for computing in cutting-edge scientific fields, such as planetary simulation, drug molecule design, gene analysis, etc.

[0037] The "intelligent computing center" described in the present invention refers to: a facility that provides the required computing power, data, and algorithms mainly for artificial intelligence applications (such as scenarios like artificial intelligence deep learning model development, model training, and model inference) by using large-scale heterogeneous computing power resources, including general computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from the underlying computing power to the top-level application enablement.

[0038] The "intelligent computing center" described in the present invention includes but is not limited to the "intelligent computing center".

[0039] The "intelligent computing center" described in the present invention, that is, the artificial intelligence computing center, is a type of computing power infrastructure that provides computing power services, data services, and algorithm services required for artificial intelligence applications based on artificial intelligence theory and using an artificial intelligence computing architecture.

[0040] The "computing power center" described in the present invention refers to a facility mainly composed of infrastructure such as wind, fire, water, and electricity, and IT software and hardware devices, which has computing power, carrying capacity, and storage capacity, including general data centers, intelligent computing centers, supercomputing centers, etc.

[0041] The "supercomputing center" described in the present invention refers to a supercomputing data center, which is a data center based on supercomputers or large-scale computing clusters, and can provide functions such as large-scale computing, storage, and network services, and is widely used in application scenarios such as aerospace, national defense, oil exploration, climate modeling, and genome sequencing.

[0042] The "computing power resources" described in the present invention refer to technologies and facilities with information computing, transmission, storage, and application capabilities required for the development of the digital society, including but not limited to computing resources such as CPUs and GPUs, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and support and guarantee resources such as wind, fire, water, and electricity.

[0043] The "large model" described in the present invention includes but not limited to "large language model" and "multimodal large model".

[0044] The "large language model" described in the present invention refers to a large language model (LLM), which is a language model with a relatively large number of parameters, aiming to understand and generate human language, and is trained through a large amount of text data, and can perform a wide range of tasks including text summarization, translation, sentiment analysis, etc.

[0045] The "multimodal large model" (Multimodal Large Models) described in the present invention refers to a model that jointly trains multimodal information such as text, images, videos, and audio, including but not limited to multimodal large language models.

[0046] The "inclusive computing power" described in the present invention refers to providing appropriate and effective computing power services to all social strata and groups with computing power service needs at an affordable cost based on the requirements of equal opportunity and the principle of commercial sustainability.

[0047] To solve the problem that existing intelligent computing centers need to load large models from remote databases, with high economic costs, low loading efficiency, and difficulty in realizing the wide application of inclusive computing power, please refer to Figure 1 , the present invention provides a large model hierarchical caching and loading method for an inclusive computing power intelligent computing center, and the method includes: Step S1: Real-time statistics of the usage times of the model layers stored in the remote database, where the model layers are obtained by disassembling a large model, and each large model stored in the remote database is disassembled into multiple model layers for storage; Step S2: When the usage times of the model layer exceed the first preset threshold, copy the model layer from the remote database to the first storage device of the local computing node; Step S3: When the usage times of the model layer exceed the second preset threshold, determine whether the model layer is a public model layer or a private model layer. If the model layer is a public model layer, move the model layer from the first storage device to the second storage device of the local computing node, where the second preset threshold is greater than the first preset threshold, the access speed of the second storage device is higher than that of the first storage device, the public model layer refers to the model layer shared by multiple tenants, and the private model layer refers to the model layer that only belongs to one tenant.

[0048] It should be noted that the tenant in the present invention refers to the tenant who leases the computing power resource storage space of the intelligent computing center. There can be multiple tenants, and the user is the user served by the tenant. The user can store the large model in the computing power resource storage space of the tenant serving it, and one tenant can serve multiple users.

[0049] In the present invention, the large model of the user is disassembled to obtain multiple model layers, which are stored in the remote database. The usage times of the model layers stored in the remote database are statistically analyzed in real time. According to the usage times of the model layers, the model layers are cached in the storage device of the local computing node. Therefore, when the model layer needs to be loaded, if the model layer is stored in the storage device of the local computing node, it can be directly loaded from the storage device of the local computing node into the memory, which can effectively save the network traffic used for loading the large model, and effectively improve the loading speed of the large model, greatly reducing the economic cost and promoting the wide application of inclusive computing power.

[0050] Please refer to Table 1. Table 1 shows the caching rules of the model layers of the present invention.

[0051] Table 1 Number Model layering Hardware computing power resources Whether local storage 1 Public high-frequency usage layer High-speed storage (nvme) Yes 2 Public low-frequency usage layer Ordinary storage (disk) Yes 3 Private high-frequency usage layer Ordinary storage (disk) Yes 4 Private low-frequency usage layer Remote database No In Table 1, the public high-frequency usage layer refers to the public model layer whose usage times exceed the second preset threshold. This type of model layer can be cached in the second storage device of the local computing node, and the second storage device is a local high-speed storage device, such as a solid-state drive (nvme), etc.

[0052] The public low-frequency usage layer refers to the public model layer with the usage times exceeding the first preset threshold, which can be stored in the first storage device of the local computing node. The first storage device is a local ordinary storage device, such as a disk, etc.

[0053] The private high-frequency usage layer refers to the private model layer with the usage times exceeding the first preset threshold, which can be stored in the first storage device of the local computing node. The first storage device is a local ordinary storage device, such as a disk, etc.

[0054] The private low-frequency usage refers to the private model layer with the usage times lower than the first preset threshold, which is stored in the remote database.

[0055] In addition, the public model layer with the usage times lower than the first preset threshold is also stored in the remote database.

[0056] In the present invention, optionally, the step S2 includes: Step S21: When the usage times of the model layer exceed the first preset threshold, copy the model layer from the remote database to the first storage device of all local computing nodes of the intelligent computing center; wherein, for the local computing node where the user currently needs to read the model layer, copy the model layer from the remote database to the first storage device of the local computing node; for other local computing nodes of the intelligent computing center, monitor the usage ratio of the network bandwidth between the other local computing nodes and the remote database, and when the usage ratio of the network bandwidth is less than the preset bandwidth usage ratio threshold, copy the model layer from the remote database to the first storage device of the other local computing nodes.

[0057] That is to say, for the local computing node where the user currently needs to read the model layer, since the model layer needs to be used immediately, it is necessary to immediately copy the model layer from the remote database to the first storage device of the local computing node. For other local computing nodes, the model layer can be copied from the remote database to the first storage device of the other local computing nodes when the network is idle.

[0058] The preset bandwidth usage ratio threshold, such as 30% etc., can be flexibly set according to needs.

[0059] In the present invention, optionally, the step S3 includes: Step S31: When the usage times of the model layer exceed a second preset threshold, determine whether the model layer is a public model layer or a private model layer. If the model layer is a public model layer, for each local computing power node of the intelligent computing center, move the model layer from the first storage device of the local computing power node to the second storage device of the local computing power node.

[0060] In the present invention, since the model layer is cached on all local computing power nodes, when a user needs to read the model layer, it can be loaded from any local computing power node, thereby effectively improving the loading speed of the large model.

[0061] In the present invention, optionally, if the large model is a model obtained by fine-tuning a basic model, the large model includes two model layers, and the two model layers include: a basic model parameter layer and a fine-tuning parameter layer.

[0062] The basic model refers to a pre-trained large model, such as GPT or BERT, etc.

[0063] Large model fine-tuning refers to a technical process of secondary training based on a pre-trained large model with data in a specific domain or task to make the model adapt to a specific application scenario.

[0064] In the present invention, the large model can be a large model using LoRA fine-tuning (Low-Rank Adaptation). LoRA fine-tuning is an efficient model fine-tuning technology, and its core idea is to introduce low-rank matrices on the basis of the pre-trained model to reduce the number of parameters required for fine-tuning, thereby improving the training efficiency and avoiding overfitting.

[0065] In the present invention, optionally, the large model includes multiple model layers disassembled based on the model structure.

[0066] For example, assuming that the large model is a Transformer model architecture, the large model can be divided into two model layers: an encoding component and a decoding component, or, more specifically, divided into the self-attention layer and the feed-forward network of the encoding component, the self-attention layer, the attention layer and the feed-forward network of the decoding component, a total of 5 model layers.

[0067] In the present invention, optionally, the method for caching and loading the large model layer of the intelligent computing center for inclusive computing power further includes: Step S4: Monitor the storage capacity of the second storage device in real time. If the storage capacity of the second storage device is higher than the first preset capacity ratio, compare the usage times of each public model layer stored in the second storage device, and move the public model layer with the lowest usage times to the first storage device for storage.

[0068] The first preset capacity ratio can be, for example, 70% or 80%, etc., and can be flexibly set according to specific circumstances.

[0069] In the present invention, optionally, the large model hierarchical caching and loading method for the inclusive computing power intelligent computing center further includes: Step S5: Monitor the storage capacity of the first storage device in real time. If the storage capacity of the first storage device is higher than the second preset capacity ratio, determine whether the private model hierarchy is stored in the first storage device. If the private model hierarchy is stored in the first storage device, compare the usage times of each private model hierarchy stored in the first storage device, and delete the private model hierarchy with the lowest usage time; if the private model hierarchy is not stored in the first storage device, compare the usage times of each public model hierarchy stored in the first storage device, and delete the public model hierarchy with the lowest usage time.

[0070] The second preset capacity ratio can be, for example, 70% or 80%, etc., and can be flexibly set according to specific circumstances.

[0071] In the present invention, by monitoring the storage capacity of the local storage device in real time, the local storage device can be cleaned in a timely manner, improving the utilization rate of the local storage device.

[0072] In the present invention, optionally, the large model hierarchical caching and loading method for the inclusive computing power intelligent computing center further includes: Step S6: In response to an instruction triggered by the user to read the large model, obtain information on all model hierarchies of the large model; The information on the model hierarchy may include, for example, information that can uniquely identify the model hierarchy, such as the index value of the model hierarchy or the name of the model hierarchy. The information on the model hierarchy may also include the storage address or type of the model hierarchy. According to the storage address or type of the model hierarchy, it can be determined whether the model hierarchy is a public model hierarchy or a private model hierarchy.

[0073] Step S7: According to the information on the model hierarchy, determine whether the model hierarchy is a public model hierarchy or a private model hierarchy; Step S8: If the model layer is a common model layer, query whether the model layer is stored in the second storage device. If the model layer is stored in the second storage device, load the model layer into the memory. If the model layer is not stored in the second storage device, query whether the model layer is stored in the first storage device. If the model layer is stored in the first storage device, load the model layer into the memory. If the model layer is not stored in the first storage device, pull the model layer from the remote database, save it to the first storage device, and load the model layer into the memory. Step S9: If the model layer is a private model layer, query whether the model layer is stored in the first storage device. If the model layer is stored in the first storage device, load the model layer into the memory. If the model layer is not stored in the first storage device, pull the model layer from the remote database, save it to the first storage device, and load the model layer into the memory. Step S10: Merge all the model layers of the large model loaded into the memory to obtain the large model.

[0074] In the present invention, optionally, the method for caching and loading the large model layers for the inclusive computing power intelligent computing center further includes: Step S11: Each time the model layer is used by the user of the tenant, increment the usage count of the model layer by 1.

[0075] In the present invention, optionally, the method for caching and loading the large model layers for the inclusive computing power intelligent computing center further includes: Step S12: In response to the instruction triggered by the user to save the large model, disassemble the user's large model to obtain multiple model layers. Step S13: Calculate the hash value of the content of each model layer. In the present invention, the hash value of the content of the model layer is calculated using the Hash Algorithm. The hash algorithm is a function that maps input data of any length to an output of a fixed length. The hash algorithm has the following characteristics: Determinism: The same input will always produce the same output, ensuring that the result is verifiable.

[0076] Irreversibility: It is impossible to deduce the original data from the hash value, ensuring security.

[0077] Collision resistance: The probability that different inputs produce the same output is extremely low, reducing the risk of data tampering.

[0078] Avalanche effect: A small change in the input leads to a significant difference in the output, enhancing sensitivity.

[0079] High efficiency: Fast calculation, suitable for large-scale data or real-time scenarios.

[0080] Step S14: Retrieve whether there is an index record corresponding to the hash value of the model layer in the index library. If it exists, do not save the model layer; if not, save the model layer to the remote database, and generate an index record according to the hash value of the model layer and the storage path of the model layer and store it in the index library; In the present invention, multiple index records can be stored in the index library, and each index record includes the hash value of a model layer and the storage path of this model layer.

[0081] For example, the index records stored in the index library can be as follows: index1={Hash1, path1}; index2={Hash2, path2}; …… Among them, index1={Hash1, path1} is an index record, and index1 is the index value of the index record.

[0082] In the present invention, optionally, the hash value can also be used as the index value, and the index records stored in the index library can be as follows: index1={path1}; index2={path2}; …… In the present invention, it should be noted that when the hash value is not used as the index value, but the index value is set separately, each hash value uniquely corresponds to an index value.

[0083] In the present invention, optionally, when saving the model layer, it can be stored in the form of a file. Optionally, the index record can also include the file size (size) of the model layer.

[0084] In the present invention, optionally, the index record can also include the tenant sharing times (or tenant reference times, count) of the model layer, and the tenant sharing times are used to indicate how many tenants share this model layer.

[0085] For example, the index records stored in the index library can be as follows: index1={path1,size1,count1}; index2={path2,size2,count2}; …… In the present invention, optionally, the index record may further include the usage times of the model layer (also referred to as the user usage times), and the usage times are used to indicate how many users use the model layer. The usage times in the present invention may include two types of usage times. The first is how many users store it, and the second is how many users read it. The usage times mentioned in the above embodiments refer to how many users read it.

[0086] In the present invention, when the index record includes the file size, the tenant sharing times, and / or the usage times, an index record is generated according to the hash value of the model layer, the storage path of the model layer, and the file size, tenant sharing times, and / or usage times of the model layer, and the generated index record is stored in the index library.

[0087] Step S15: Record the index values of all model layers of the large model.

[0088] Optionally, the index value of the model layer is the hash value of the model layer.

[0089] For example, a certain large model can be recorded as follows: Model1: {index1, index2}. That is, Model1 is disassembled into two model layers, where the index values of the two model layers are index1 and index2 respectively.

[0090] Optionally, when recording the index values of all model layers of the large model, the storage path and / or the file size of the model layer can also be recorded simultaneously.

[0091] For example, a certain large model can be recorded as follows: Model1 { index1= / c: / publish / model_level1,50G index2= / c: / private / model_level2,1G } Among them, / c: / publish / model_level1 is the storage path of the model layer, and 50G is the file size of the model layer.

[0092] In the present invention, when a user needs to store a large model in an intelligent computing center, the large model of the user is disassembled to obtain multiple model layers. The hash value of the content of each model layer is calculated, and it is retrieved whether there is an index record corresponding to the hash value of the model layer in the index library. If it exists, it means that the content of this model layer has been stored before, so there is no need to store it repeatedly. If it does not exist, it means that the content of this model layer has not been stored before, then this model layer is stored, and an index record is generated according to the hash value of the model layer and the storage path of the model layer and stored in the index library. When the content of the model layer of the large model has been stored before, there is no need to store it repeatedly, that is, there is no need to store the large model in full, so that the storage space can be effectively saved, the waste of storage and network traffic of computing resources can be reduced, the computing resources of the intelligent computing center can be optimized, the waste of computing resources can be greatly reduced, the economic cost can be greatly reduced, and the wide application of inclusive computing power can be promoted.

[0093] In the present invention, optionally, the step S14 includes: Step S141: Use the hash value of the model layer as the index value, generate an index record according to the index value of the model layer and the storage path of the model layer, and store the generated index record in the index library.

[0094] Using the hash value of the model layer directly as the index value can eliminate the need to introduce an index value separately and make the retrieval of the index library more convenient.

[0095] In the present invention, optionally, the step S14 further includes: Step S142: If there is no index record corresponding to the hash value of the model layer in the index library, save the model layer as a private model layer in the private database of the remote database, generate an index record corresponding to the hash value of the model layer and store it in the index library, and record the tenant to which the user belongs in the index record; and / or Step S143: If there is an index record corresponding to the hash value of the model layer in the index library, determine whether the tenant to which the user belongs is recorded in the index record; Step S144: If the tenant to which the user belongs is not recorded in the index record, add 1 to the number of sharing times of the tenant in the index record and record the tenant to which the user belongs in the index record; determine whether the model layer is saved in the private database. If the model layer is saved in the private database, move the model layer as a public model layer to the public database of the remote database; Step S145: If the tenant to which the user belongs has been recorded in the index record, there is no need to add 1 to the number of sharing times of the tenant in the index record.

[0096] In the present invention, during the above judgment process, if it is determined that the tenant to which the user belongs is recorded in the index record, it is not necessary to increment the tenant sharing count in the index record by 1.

[0097] In the present invention, during the above judgment process, if it is determined that the model layer is stored in the public database, it is not necessary to perform further storage processing.

[0098] In the present invention, optionally, each tenant has its own private database, and the public database is shared by multiple tenants.

[0099] In the present invention, the intelligent computing center provides a private database and a public database. If there is no index record corresponding to the hash value of a certain model layer in the index library, it indicates that the model layer has not been stored before, and the model layer can be stored in the private database.

[0100] If there is an index record corresponding to the hash value of a certain model layer in the index library, it indicates that the model layer has been stored before. At this time, it is also necessary to determine whether the tenant that currently needs to store the model layer is recorded in the index record. If the tenant has been recorded in the index record, there is no need to increase the tenant sharing count of the index record.

[0101] If the tenant is not recorded in the index record, the tenant needs to be recorded in the index record, and the tenant sharing count of the index record is incremented by 1. Then, it is also necessary to determine whether the model layer is currently stored in the private database. If so, the model layer needs to be moved to the public database. If the model layer is currently already stored in the public database, there is no need to save it again.

[0102] In the present invention, optionally, the method for caching and loading the large model layer of the intelligent computing center for inclusive computing power further includes: Step S16: Determine the storage cost for each tenant according to the model type, tenant sharing count, and / or file size of the stored model layer; Among them, for the model layer shared by multiple tenants, according to the tenant sharing count of the model layer, the storage cost of the model layer is evenly divided among the multiple tenants; or, for the model layer shared by multiple tenants, if the model layer is a basic model, no storage cost is charged to the tenants for the basic model; and / or For the model layer private to the tenant, calculate the storage cost of the model layer according to the file size of the model layer.

[0103] In the present invention, optionally, the model layer private to the tenant refers to the model layer stored in the private database.

[0104] In the present invention, the large model storage platform provided by the intelligent computing center can have multiple tenants, and each tenant serves multiple users. For each tenant, the storage cost of the large model stored by it can be calculated.

[0105] In some examples, it is not excluded to calculate the storage cost for each user.

[0106] In the present invention, optionally, the step S15 includes: Step S151: Use the hash value of the model layer as the index value, and record the index values of all model layers of the large model.

[0107] In the present invention, each large model can have a unique identifier, such as Model1. The intelligent computing center can separately provide an identifier library for recording the index values of all model layers of the large model.

[0108] For example, the record corresponding to a certain large model is as follows: Model1 {index1= / c: / publish / model_level1,50G index2= / c: / private / model_level2,1G } Please refer to Figure 2 , the present invention also provides a large model hierarchical cache loading device 100 for an inclusive computing power intelligent computing center, including: A statistics module 101, configured to statistically count the usage times of model layers stored in a remote database in real time, where the model layers are obtained by disassembling a large model, and each large model stored in the remote database is disassembled into multiple model layers for storage; A first cache module 102, configured to copy the model layer from the remote database to a first storage device of a local computing power node when the usage times of the model layer exceed a first preset threshold; A second cache module 103, configured to determine whether the model layer is a public model layer or a private model layer when the usage times of the model layer exceed a second preset threshold. If the model layer is a public model layer, move the model layer from the first storage device to a second storage device of the local computing power node, where the second preset threshold is greater than the first preset threshold, the access speed of the second storage device is higher than that of the first storage device, the public model layer refers to a model layer shared by multiple tenants, and the private model layer refers to a model layer that belongs to only one tenant.

[0109] In the present invention, the user's large model is disassembled to obtain multiple model layers, which are stored in a remote database. The usage times of the model layers stored in the remote database are statistically calculated in real time. According to the usage times of the model layers, the model layers are cached to the storage device of the local computing nodes, so that when the model layers need to be loaded, if the model layers are stored in the storage device of the local computing nodes, they can be directly loaded from the storage device of the local computing nodes into the memory, which can effectively save the network traffic used for loading the large model, and effectively improve the loading speed of the large model, greatly reducing the economic cost and promoting the wide application of inclusive computing power.

[0110] In the present invention, optionally, the first caching module 102 is configured to, when the usage times of the model layer exceed a first preset threshold, copy the model layer from the remote database to the first storage devices of all local computing nodes of the intelligent computing center. For the local computing node where the user currently needs to read the model layer, copy the model layer from the remote database to the first storage device of the local computing node; for other local computing nodes of the intelligent computing center, monitor the usage ratio of the network bandwidth between the other local computing nodes and the remote database, and when the usage ratio of the network bandwidth is less than a preset bandwidth usage ratio threshold, copy the model layer from the remote database to the first storage devices of the other local computing nodes.

[0111] In the present invention, optionally, the second caching module 103 is configured to, when the usage times of the model layer exceed a second preset threshold, determine whether the model layer is a public model layer or a private model layer. If the model layer is a public model layer, for each local computing node of the intelligent computing center, move the model layer from the first storage device of the local computing node to the second storage device of the local computing node.

[0112] In the present invention, optionally, if the large model is a model obtained by fine-tuning a basic model, the large model includes two model layers, and the two model layers include: a basic model parameter layer and a fine-tuning parameter layer.

[0113] In the present invention, optionally, the large model includes multiple model layers disassembled based on the model structure.

[0114] In the present invention, optionally, the large model layer caching and loading device 100 for the inclusive computing power intelligent computing center further includes: The first monitoring module is used to monitor the storage capacity of the second storage device in real time. If the storage capacity of the second storage device is higher than the first preset capacity ratio, compare the usage times of each public model layer stored in the second storage device, and move the public model layer with the lowest usage times to the first storage device for storage.

[0115] In the present invention, optionally, the large model layer caching and loading device 100 for the inclusive computing power intelligent computing center further includes: The second monitoring module is used to monitor the storage capacity of the first storage device in real time. If the storage capacity of the first storage device is higher than the second preset capacity ratio, determine whether the private model layer is stored in the first storage device. If the private model layer is stored in the first storage device, compare the usage times of each private model layer stored in the first storage device, and delete the private model layer with the lowest usage times; if the private model layer is not stored in the first storage device, compare the usage times of each public model layer stored in the first storage device, and delete the public model layer with the lowest usage times.

[0116] In the present invention, optionally, the large model layer caching and loading device 100 for the inclusive computing power intelligent computing center further includes: The acquisition module is used to acquire information of all model layers of the large model in response to a command triggered by the user to read the large model. The judgment module is used to judge whether the model layer is a public model layer or a private model layer according to the information of the model layer. The first query module is used to, if the model layer is a public model layer, query whether the model layer is stored in the second storage device. If the model layer is stored in the second storage device, load the model layer into the memory; if the model layer is not stored in the second storage device, query whether the model layer is stored in the first storage device; if the model layer is stored in the first storage device, load the model layer into the memory, if the model layer is not stored in the first storage device, pull the model layer from the remote database and save it to the first storage device, and load the model layer into the memory. The second query module is used to, if the model layer is a private model layer, query whether the model layer is stored in the first storage device; if the model layer is stored in the first storage device, load the model layer into the memory, if the model layer is not stored in the first storage device, pull the model layer from the remote database and save it to the first storage device, and load the model layer into the memory. A merging module, configured to merge all the model layers of the large model loaded into the memory to obtain the large model.

[0117] In the present invention, optionally, the large model hierarchical caching and loading device 100 for the inclusive computing power intelligent computing center further includes: A counting module, configured to increment the usage count of the model layer by 1 each time the model layer is used by a user of the tenant.

[0118] In the present invention, optionally, the large model hierarchical caching and loading device 100 for the inclusive computing power intelligent computing center further includes: A disassembling module, configured to disassemble the large model of the user in response to an instruction to save the large model triggered by the user, to obtain a plurality of model layers; A hash value calculation module, configured to calculate the hash value of the content of each model layer; A storage module, configured to retrieve whether there is an index record corresponding to the hash value of the model layer in the index library. If so, the model layer is not saved; if not, the model layer is saved to the remote database, and an index record is generated according to the hash value of the model layer and the storage path of the model layer and stored in the index library; A recording module, configured to record the index values of all the model layers of the large model.

[0119] Optionally, the storage module is configured to use the hash value of the model layer as the index value, generate an index record according to the index value of the model layer and the storage path of the model layer, and store the generated index record in the index library.

[0120] Optionally, the storage module is configured to, if there is no index record corresponding to the hash value of the model layer in the index library, save the model layer as a private model layer to the private database of the remote database, generate an index record corresponding to the hash value of the model layer and store it in the index library, and the index record records the tenant to which the user belongs.

[0121] Optionally, the storage module is configured to, if there is an index record corresponding to the hash value of the model hierarchy in the index library, determine whether the tenant to which the user belongs is recorded in the index record; if the tenant to which the user belongs is not recorded in the index record, increment the tenant sharing count in the index record by 1, and record the tenant to which the user belongs in the index record; determine whether the model hierarchy is stored in the private database, and if the model hierarchy is stored in the private database, move the model hierarchy as a public model hierarchy to the public database of the remote database; if the tenant to which the user belongs is already recorded in the index record, there is no need to increment the tenant sharing count in the index record.

[0122] Please refer to Figure 3 , the present invention also provides a computing power device 200, including a processor 201, a memory 202, and a computer program stored on the memory 202 and executable on the processor 201. When the computer program is executed by the processor 201, it implements each process of the above-mentioned method embodiment of the large model hierarchical cache loading method for the inclusive computing power intelligent computing center, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0123] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements each process of the above-mentioned method embodiment of the large model hierarchical cache loading method for the inclusive computing power intelligent computing center, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here. Among them, the computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.

[0124] The present invention also provides a computer program product, including computer instructions, which when executed by a processor implement the above Figure 1 each process of the method embodiment of the large model hierarchical cache loading method for the inclusive computing power intelligent computing center shown, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0125] It should be noted that in this article, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including that element.

[0126] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present invention.

[0127] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the purpose of the present invention and the scope protected by the claims, and all of them belong to the protection scope of the present invention.

Claims

1. A large model hierarchical cache loading method for an inclusive computing power intelligent computing center, characterized in that, Including: Step S1: Statistically count the usage times of the model layers stored in the remote database in real time. The model layers are obtained by disassembling a large model, and each large model stored in the remote database is disassembled into multiple model layers for storage; Step S2: When the usage times of the model layer exceed a first preset threshold, copy the model layer from the remote database to the first storage device of the local computing node; Step S3: When the usage times of the model layer exceed a second preset threshold, determine whether the model layer is a public model layer or a private model layer. If the model layer is a public model layer, move the model layer from the first storage device to the second storage device of the local computing node. The second preset threshold is greater than the first preset threshold, the access speed of the second storage device is higher than that of the first storage device. The public model layer refers to a model layer shared by multiple tenants, and the private model layer refers to a model layer that belongs to only one tenant.

2. The method according to claim 1, wherein: The step S2 includes: Step S21: When the usage times of the model layer exceed a first preset threshold, copy the model layer from the remote database to the first storage device of all local computing nodes of the intelligent computing center; for the local computing node where the user currently needs to read the model layer, copy the model layer from the remote database to the first storage device of the local computing node; for other local computing nodes of the intelligent computing center, monitor the usage ratio of the network bandwidth between the other local computing nodes and the remote database. When the usage ratio of the network bandwidth is less than the preset bandwidth usage ratio threshold, copy the model layer from the remote database to the first storage device of the other local computing nodes; And / or, the step S3 includes: Step S31: When the usage times of the model layer exceed a second preset threshold, determine whether the model layer is a public model layer or a private model layer. If the model layer is a public model layer, for each local computing node of the intelligent computing center, move the model layer from the first storage device of the local computing node to the second storage device of the local computing node.

3. The method according to claim 1, wherein If the large model is a model obtained by fine-tuning a basic model, the large model includes two model layers, and the two model layers include: a basic model parameter layer and a fine-tuning parameter layer; And / or, the large model includes multiple model layers disassembled based on the model structure.

4. The method according to claim 1, wherein It further includes: Step S4: Monitor the storage capacity of the second storage device in real time. If the storage capacity of the second storage device is higher than the first preset capacity ratio, compare the usage times of each public model layer stored in the second storage device, and move the public model layer with the lowest usage times to the first storage device for storage.

5. The method according to claim 1, wherein It further includes: Step S5: Monitor the storage capacity of the first storage device in real time. If the storage capacity of the first storage device is higher than the second preset capacity ratio, determine whether the private model layers are stored in the first storage device. If the private model layers are stored in the first storage device, compare the usage times of each of the private model layers stored in the first storage device, and delete the private model layer with the lowest usage time; if the private model layers are not stored in the first storage device, compare the usage times of each of the public model layers stored in the first storage device, and delete the public model layer with the lowest usage time.

6. The method according to claim 1, characterized in that Further included are: Step S6: In response to an instruction triggered by the user to read the large model, obtain information on all model layers of the large model; Step S7: Based on the information of the model layers, determine whether the model layer is a public model layer or a private model layer; Step S8: If the model layer is a public model layer, query whether the model layer is stored in the second storage device. If the model layer is stored in the second storage device, load the model layer into the memory; if the model layer is not stored in the second storage device, query whether the model layer is stored in the first storage device; If the model layer is stored in the first storage device, load the model layer into the memory. If the model layer is not stored in the first storage device, pull the model layer from the remote database and save it to the first storage device, and load the model layer into the memory; Step S9: If the model layer is a private model layer, query whether the model layer is stored in the first storage device; If the model layer is stored in the first storage device, load the model layer into the memory. If the model layer is not stored in the first storage device, pull the model layer from the remote database and save it to the first storage device, and load the model layer into the memory; Step S10: Merge all the model layers of the large model loaded into the memory to obtain the large model.

7. The method according to claim 6, characterized in that, Further included are: Step S11: Each time the model layer is used by the user of the tenant, increment the usage times of the model layer by 1.

8. The method according to claim 1, wherein Before the said Step S1, further included is: Step S12: In response to an instruction triggered by the user to save the large model, disassemble the user's large model to obtain multiple model layers; Step S13: Calculate the hash value of the content of each model layer; Step S14: Retrieve whether there is an index record corresponding to the hash value of the model layer in the index library. If it exists, do not save the model layer; if it does not exist, save the model layer to the remote database, and generate an index record based on the hash value of the model layer and the storage path of the model layer and store it in the index library; Step S15: Record the index values of all model layers of the large model.

9. The method according to claim 8, wherein The said Step S14 includes: Step S141: Use the hash value of the model layer as the index value, generate an index record based on the index value of the model layer and the storage path of the model layer, and store the generated index record in the index library.

10. The method according to claim 8, wherein The step S14 further includes: Step S142: If there is no index record corresponding to the hash value of the model layer in the index library, save the model layer as a private model layer in the private database of the remote database, generate an index record corresponding to the hash value of the model layer and store it in the index library, and the tenant to which the user belongs is recorded in the index record; and / or Step S143: If there is an index record corresponding to the hash value of the model layer in the index library, determine whether the tenant to which the user belongs is recorded in the index record; Step S144: If the tenant to which the user belongs is not recorded in the index record, increment the tenant sharing count in the index record by 1 and record the tenant to which the user belongs in the index record; determine whether the model layer is saved in the private database, and if the model layer is saved in the private database, move the model layer as a public model layer to the public database of the remote database; Step S145: If the tenant to which the user belongs has been recorded in the index record, there is no need to increment the tenant sharing count in the index record.

Citation Information

Patent Citations

  • Cross-data center computing power resource scheduling method and device based on container

    CN118916164A

  • Model loading method and system

    CN119066851A

  • Hierarchical thresholds-based virtual machine configuration

    US20140196030A1

  • Nested group hierarchies for analytics applications

    US20210390120A1