Large model deployment method and device, computer equipment and storage medium

By establishing containers and generating image files in large-scale model deployment, combined with grouping and quantification methods, the problems of excessive bandwidth and storage space in traditional deployment processes are solved, efficient and secure cloud service operation is achieved, and the management and deployment process of large models is simplified.

CN120743293APending Publication Date: 2025-10-03CHINA PING AN LIFE INSURANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510823949.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

The traditional large-scale AI model deployment process faces problems such as excessive transmission bandwidth, excessive storage space, and difficult deployment and operation and maintenance. These problems limit the rapid iteration and growth of cloud applications, and pose a threat to the security, stability, and efficiency of AI services and cloud services in the fields of financial technology, healthcare, and elderly care.

Method used

By establishing a container, encapsulating the large model in binary format, generating an image file, associating the model file with the image file, and transferring it to the cloud storage device, container technology is used to simplify the deployment process, reduce bandwidth and computing power resource consumption, and combine grouping and quantization methods to compress storage space.

Benefits of technology

It reduces the transmission bandwidth and storage resource consumption of large models, improves the management and use efficiency of the system, ensures the safe, stable and smooth operation of cloud services, simplifies the user's self-deployment process, and facilitates the development and management of large models across multiple business lines and technology stacks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120743293A_ABST
    Figure CN120743293A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of large models, is suitable for the fields of financial science and technology and medical health care, and discloses a large model deployment method and device, computer equipment and a storage medium. The method comprises the following steps: establishing a first container, loading a target large model through the first container, and operating the target large model; packaging the target large model in a grouping mode and / or a quantification mode to obtain a model file in a binary format; deleting the target large model in the first container, and generating a mirror image file according to the first container; associating the model file with the mirror image file to obtain associated information data; and transmitting the associated information data, the model file and the mirror image file to cloud storage equipment. According to the large model deployment method, consumption in the aspects of transmission bandwidth, storage resources, deployment processes, computing power operation and maintenance and the like of the large model can be reduced, and the efficiency of a management and use system of the large model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of large model technology, and in particular to a large model deployment method, apparatus, computer equipment, and storage medium. Background Art

[0002] With the widespread application of artificial intelligence (AI) in financial technology, healthcare, and senior care, their reliance on cloud infrastructure is increasing. As a key technology supporting enterprise AI transformation, cloud technology provides powerful computing power and storage support for innovative applications such as intelligent risk control and precision marketing in financial technology, as well as remote diagnosis and personalized treatment plans in healthcare. However, traditional cloud technology solutions for deploying large AI models face numerous challenges, including high transmission bandwidth, excessive storage space, and difficult deployment and maintenance. These issues not only limit the rapid iteration and growth of cloud applications, but also pose a threat to the security, stability, and efficiency of AI and cloud services in financial technology, healthcare, and senior care. Summary of the Invention

[0003] The present application discloses a large model deployment method, device, computer equipment and storage medium, which solves the problem of low efficiency in the transmission, storage, management, deployment and use of large models in related technologies. It can reduce the consumption of transmission bandwidth, storage resources, deployment process and computing power operation and maintenance of the large model itself, improve the management and use system efficiency of the large model, reduce latency, save bandwidth and computing power resources, and ensure the safe, stable and smooth operation of cloud services.

[0004] In a first aspect, the present application provides a large model deployment method, the method comprising:

[0005] Establishing a first container, and running the target large model through the first container;

[0006] Encapsulating the target large model in a grouping manner and / or a quantized manner to obtain a model file in a binary format;

[0007] Deleting the target large model in the first container and generating an image file based on the first container;

[0008] Associating the model file with the image file to obtain associated information data, wherein the associated information data includes model version management data and image version management data, and a corresponding relationship between the model file and the image file can be obtained based on the model version management data and the image version management data;

[0009] The associated information data, the model file, and the image file are transmitted to a cloud storage device.

[0010] In a second aspect, the present application provides a large model deployment device, comprising:

[0011] A container testing module, configured to establish a first container and run a target large model through the first container;

[0012] A model packaging module is used to package the target large model to obtain a model file;

[0013] an image creation module, configured to delete the target macro model of the first container and generate an image file based on the first container;

[0014] A file association module, associating the model file with the image file to obtain association information data;

[0015] The storage transmission module transmits the associated information data, the model file and the image file to a cloud storage device.

[0016] In a third aspect, the present application provides a computer device, comprising a processor, a memory, and a computer program stored in the memory and executable by the processor, wherein when the computer program is executed by the processor, it implements the method provided in any embodiment of the present application.

[0017] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the processor implements a method as provided in any embodiment of the present application.

[0018] The large model deployment method proposed in this application is encapsulated based on container technology, which can better integrate the processes of management, development, operation and maintenance, and compression of the target large model, avoid overly complicated operations in related technologies, and reduce the implementation threshold and computing power consumption; by encapsulating in a grouping and / or quantization manner and saving the model file in a binary format, the storage space occupied by the model file can be further compressed, and the bandwidth resources required for subsequent transmission to the cloud storage device can be saved; by associating model files and image files with associated information data, it not only greatly simplifies the user's possible self-deployment process, but also facilitates large enterprises and institutions to develop and manage large models of multiple business lines and multiple technology stacks.

[0019] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0021] Figure 1 This is a schematic flow chart of the steps of a large model deployment method provided in one embodiment of the present application;

[0022] Figure 2 This is a schematic flow chart of the steps of a large model deployment method provided in one embodiment of the present application;

[0023] Figure 3 This is a schematic flow chart of the steps of a large model deployment method provided in one embodiment of the present application;

[0024] Figure 4 This is a schematic flow chart of the steps of a large model deployment method provided in one embodiment of the present application;

[0025] Figure 5 This is a structural diagram of a large model deployment device provided in an embodiment of the present application;

[0026] Figure 6 This is a schematic block diagram of the structure of a computer device provided in an embodiment of the present application.

[0027] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. DETAILED DESCRIPTION

[0028] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0029] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, combined, or partially merged, so the actual execution order may vary depending on the actual situation.

[0030] It should be understood that the terms used in this specification are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in this specification and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0031] It should be understood that, in order to facilitate a clear description of the technical solutions of the embodiments of the present application, in the embodiments of the present application, words such as "first" and "second" are used to distinguish between identical or similar items with substantially the same functions and effects. For example, the first data and the second data are merely used to distinguish different data and do not limit their order. Those skilled in the art will understand that words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not necessarily limit differences.

[0032] It should be further understood that the term “and / or” used in this specification and the appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0033] The following describes some embodiments of the present application in detail with reference to the accompanying drawings. In the absence of conflict, the following embodiments and features therein may be combined with each other.

[0034] With the rapid development of cloud computing, cloud storage, and big model technologies, cloud systems have gradually become the key infrastructure for enterprises to enable technological innovation and market expansion, and the research and development of big model technologies has gradually become a hot spot for corporate technological competition. Especially in the fields of financial technology and medical health and elderly care, cloud infrastructure and big model deployment, management, and operation and maintenance systems have gradually become important enabling tools for the industry's informatization and intelligence, but have also brought new problems and challenges.

[0035] For example, in the field of financial technology, especially in key scenarios such as open banking, high-frequency trading, and compliance monitoring, large model application nodes such as risk control models, data analysis centers, model training centers, terminal models, and mini-program-based large model services often require frequent adjustments and updates to adapt to various data sources and server-based operation and maintenance environments. This is especially true for high-frequency trading and risk control models, which require rapid response to market changes and analysis and prediction of real-time data. The technical disadvantage of traditional deployment methods, which strongly couple models with the operating environment, means that each model or environment update requires redeployment and adjustment, consuming significant time and resources for R&D and operation and maintenance users. Furthermore, frequent model updates and high concurrency requests place extremely high demands on the computing power and bandwidth of cloud systems, making traditional deployment methods difficult to meet the needs of dynamic scaling.

[0036] For example, in the healthcare and elderly care sectors, particularly in telemedicine, health management, electronic medical records, medical image storage and transmission, and intelligent medical guidance, the rapid development of cloud technology has given doctors, patients, and hospital administrators the opportunity to apply a wide range of medical big models, improving the efficiency of research, management, diagnosis, and treatment in the healthcare sector. However, because models for different diseases and departments require the integration of their own databases and domain knowledge to generate treatment recommendations tailored to the patient's specific situation, they must be adaptable to the diverse equipment and operating environments of different medical institutions and medical big model service providers to meet the requirements of sensitive data privacy management. For medical institutions, high R&D personnel costs and diverse deployment environments have always been pain points and difficulties in deploying big model services. How to abstract the device environment in a lightweight manner, simplify model deployment for medical institutions, and achieve version control and decouple model-environment management has become a pressing technical challenge.

[0037] To solve some of the above problems, the present application proposes a large model deployment method. Figure 1 , Figure 1 This is a schematic flow chart of the steps of the large model deployment method provided by an embodiment of the present application. Figure 1 As shown, the large model deployment method may specifically include steps S101 to S104.

[0038] S101: Establish a first container and run a target large model through the first container.

[0039] It should be understood that the target large model must be executable within the first container and is not a static "model file." The process of running the target large model can include loading and launching the target large model via a network or storage medium, or building and training the target large model locally. This is not specifically limited here.

[0040] In some embodiments, while installing the necessary dependencies on top of the base image, you can also install the necessary deep learning frameworks based on the requirements of the target large model. After the large model is launched, you can monitor the container logs in real time to ensure the model is running properly, and stop and clean up the container when necessary.

[0041] In some embodiments, the CPU, GPU, and memory resources can be queried through command-line tools, task managers, or system information tools to ensure that there are sufficient hardware resources. For example, the detailed information of the CPU, including the number of cores, number of threads, CPU frequency, etc., can be queried through the lscpu command. As another example, the nvidia-smi command can be used to query the usage of the GPU, including GPU utilization, video memory usage, etc. As another example, the official documentation of the target large model can be consulted to understand its recommended hardware configuration. For example, some large models may require at least 16GB of GPU video memory and 32GB of system memory. If there is a model training log, the usage of hardware resources can also be extracted from it for reference.

[0042] Furthermore, if queries reveal that the device running the container has limited resources, solutions can be found by increasing hardware resources, optimizing the model, and implementing distributed deployment. For example, if existing hardware resources are insufficient, consider increasing the number of CPU cores, memory capacity, or GPUs. Alternatively, if increasing hardware resources is not possible, consider optimizing the model, such as using quantization techniques or model pruning, to reduce the model's demand for hardware resources. Alternatively, shard the model and deploy it across multiple devices, leveraging distributed computing to meet the model's operational needs.

[0043] You should understand that by querying hardware resources, assessing model requirements, using resource monitoring tools, and setting container resource limits, you can ensure that there are sufficient hardware resources to support the operation of large models. This not only helps improve model performance and stability, but also avoids operation failures due to insufficient resources.

[0044] In some embodiments, large models typically require a series of software, tools, scripts, and files to support their operation. These dependencies are organized into a dependency tree to ensure that the large model can function properly in a specific operating system environment. Furthermore, for ordinary users in medical institutions and fintech companies, their professional technical capabilities are relatively limited. For large model service providers, specific application interfaces can be established for the large model to enhance its functions in data processing, communication query, computing power assistance, application collaboration, human-computer interaction, and other aspects.

[0045] See also Figure 2 , Figure 2 This is a schematic flow chart of the steps of the large model deployment method provided by an embodiment of the present application. Figure 2 As shown, the large model deployment method may specifically include steps S101a to S101d, which are used to implement the above step S101.

[0046] S101a: Create a first container and run an operating system through the first container.

[0047] It should be understood that a suitable base image can be selected to establish the first container. The base image may include an operating system suitable for running the large model. The base image may not include an operating system suitable for running the large model. The first container can be established using the base image first, and then an operating system suitable for running the large model can be loaded and run in the first container.

[0048] In some embodiments, the base image can be a lightweight operating system distribution that provides a minimal, simplified operating system environment but does not include configurations suitable for the large model. Therefore, after creating the first container, you can continue to install and configure an operating system suitable for running the large model using the package management tools within the container.

[0049] In some embodiments, a basic image based on a relatively complete operating system can be selected to build a container, and then necessary operating system components and dependent libraries, such as programming language interpreters, compilers, deep neural network frameworks, etc., can be installed in the container to build an environment suitable for running large models.

[0050] It should be understood that regardless of the method chosen, the first container establishment process should ensure that the necessary hardware resources and software environment are provided for the operation of the large model. This includes but is not limited to allocating sufficient CPU cores, memory capacity, and GPU resources, and ensuring that the operating system and dependent library versions match the requirements of the large model. Through the above steps, the first container can be flexibly established and run within it an operating system suitable for the large model, thus providing a stable basic environment for the efficient operation of the large model.

[0051] S101b. According to the current situation of the first container and its operating system, a target macro model adapted to the current situation is obtained.

[0052] It should be understood that the current situation may include the inherent parameters and state parameters of the first container and its operating system. For example, it may include the hardware resources allocated to the first container, including inherent parameters such as the number of CPU cores, memory capacity, and GPU resources. Specifically, the resource usage of the current container can be obtained through a container management tool (such as Docker's docker stats command), and the operating system version of the first container and the software package management tool it supports can be confirmed. For another example, it can be confirmed whether the kernel version and system configuration of the operating system meet the operating requirements of the target large model, and whether the operating system has installed the necessary dependent libraries and tools.

[0053] In some embodiments, a target large model suitable for the current situation can be retrieved from the model repository based on the resources of the first container and the compatibility of the operating system. For example, if the first container has high-performance GPU resources, a large model version that supports GPU acceleration can be selected. It should be understood that the model repository can be a public model repository or a private model storage service, such as an internal enterprise model warehouse.

[0054] In some embodiments, it is possible to check and ensure that the downloaded target large model matches the operating system and hardware resources of the first container. For example, if the first container uses a server operating system, the model file suitable for the server operating system should be downloaded. It is also possible to verify and ensure that the target large model can run normally after loading it in the first container. For example, the model's loading speed, inference speed, and resource usage can be checked by running the model's test script or sample code. If the target large model does not run properly in the current environment, the resource allocation of the first container or the configuration of the operating system can be adjusted as needed, or the adapted target large model can be retrieved.

[0055] In some embodiments, an operating system suitable for running large models can be selected, followed by the installation of container runtime tools and necessary package management tools. It should be understood that container runtime tools can be used to manage and run containers, ensuring that applications run in an isolated environment; package management tools can be used to install, update, and manage software packages in the operating system, streamlining the software installation and management process. By properly utilizing these tools, development and operations efficiency can be significantly improved.

[0056] S101c. According to the target macro model, a dependency tree and an application interface are set through the operating system.

[0057] It should be understood that establishing an application interface for a large model within a container is a key step in ensuring that the model can be called and interacted with by external systems or users. For example, a service interface can be created so that the model can receive input and return output over the network. For example, the model's input and output interfaces can be defined through service code. The service code can be used to load the large model and provide one or more endpoints to receive external requests and return model inference results.

[0058] In some embodiments, to ensure high service availability, load balancers and container orchestration tools can be used to manage multiple service instances, enabling on-demand pull, resource scheduling, and service exposure. Through these steps, application interfaces for large models can be established within the container, ensuring that the model can receive input over the network and return inference results. This not only improves model accessibility and scalability, but also facilitates model deployment and management.

[0059] In some embodiments, the application interface of the large model can not only support efficient inference services but also provide a user interface and multimedia tools to implement prompt engineering and user interaction. Prompt engineering refers to guiding the large model to generate high-quality output results by designing and optimizing input prompts. The user interface and some multimedia tools allow users to easily interact with the model, adjust input parameters, view output results, and perform further analysis and processing.

[0060] It should be understood that the large model itself lacks an intuitive user interface, making it difficult for users to conveniently interact with the model, adjust input parameters, and view output results. The large model itself also does not include functions such as preset prompts and multimedia tools. In the use of traditional large model services, users need to manually design and optimize input prompts based on different questions and conversations. The consistency of questions and answers is poor, and the use is difficult. On this basis, for medical institutions and financial technology institutions such as banks and insurance companies that lack professional technical teams related to large models, it is difficult to develop application interfaces on their own to integrate the model with existing databases and software tools, which limits the widespread application of large models. Therefore, the AI ​​runtime, dependency libraries and standardized API interfaces can be pre-integrated in a lightweight image file built based on the container, so that specific functions can be directly implemented after rapid deployment.

[0061] In some embodiments, the application interface may include a prompt word module, a security module, and / or an interactive interface module. For example, the application interface may include a prompt word module such as a prompt engineering tool to support prompt generation and configuration of prompt strategies. Specifically, prompt generation may include rule-based prompt generation, template-based prompt generation, and machine learning-based prompt optimization. Furthermore, users can select the appropriate prompt configuration strategy based on their specific needs to improve the output quality of the model.

[0062] In some embodiments, the application interface may also include interactive interface modules such as a visual interactive interface, a speech-to-text component, an image generation component, and a data table generation component, thereby allowing users to conveniently interact with the large model. For example, users can use the interface to input parameters, adjust model operating parameters, view model output results, and perform further analysis and processing.

[0063] In some embodiments, the application interface may also include a standardized API, enabling users with sufficient technical skills to integrate the model with existing systems and tools. For example, users can invoke the services of the large model through the API and integrate the model's output into existing business processes. This configuration enables the application interface to support multiple programming languages ​​and frameworks, facilitating secondary development and expansion by developers. Users can develop customized applications based on specific needs, improving the applicability and flexibility of the large model.

[0064] S101d. Run the target large model through the first container.

[0065] It's important to understand that in modern AI applications, running large models requires a stable, efficient, and scalable environment. Traditional model deployment methods often rely on complex configuration and manual operations, which not only increases deployment difficulty but can also lead to inconsistent environments and unstable operations. Furthermore, existing deployment methods often lack granular control over the model's runtime environment, failing to fully leverage the advantages of containerization technology.

[0066] In some embodiments, necessary dependencies can be installed directly through the operating system. Furthermore, taking advantage of containerization technology, a startup script can be loaded or created to define and use various commands to run when the container starts. It should be understood that the container's working directory can also be specified through a configuration file, such as a startup script, so that the model file and startup script are located in the correct location.

[0067] S102. Encapsulate the target large model by grouping and / or quantizing to obtain a model file in binary format.

[0068] It should be understood that the binary format may include GGUF (General GPU-Optimized Unified Format), GGML (Georgi Gerganov Machine Learning), ONNX, TensorFlow Lite and other formats. Among them, because the GGUF format reduces the size of the model file by optimizing the storage structure while maintaining the integrity and accuracy of the model, it performs well in high-performance reasoning scenarios. In addition, the GGUF format is specially optimized for GPU reasoning, which can make full use of the parallel computing capabilities of the GPU and significantly improve the reasoning speed. Furthermore, the GGUF format supports a variety of data types and model structures, and is suitable for a variety of deep learning frameworks and application scenarios. It is very suitable for institutional users in the medical, health and elderly care fields with diverse and complex hardware operating environments, as well as institutional users in the financial technology field such as open banking, insurance, and securities companies with diverse and specific privacy regulatory needs.

[0069] In some embodiments, the target large model can be quantized, and the weights and activation values ​​of the model can be converted from floating point numbers (such as 32-bit floating point numbers) to low-precision formats (such as 8-bit integers). This significantly reduces the storage space and computational complexity of the model while maintaining the performance of the model. The target large model can also be grouped, and the large model can be divided into multiple small model files, each file containing a portion of the model's parameters and computing tasks. The grouping algorithm can also be mixed with the quantization algorithm for more efficient packaging. For example, the quantized model parameters can be stored in slices, and each slice file contains part of the quantized parameters and computing tasks.

[0070] It's important to understand that data security and privacy protection are paramount in the healthcare and fintech sectors. Medical institutions and financial institutions typically have stringent security measures in place to protect user data and models. However, traditional model deployment methods often require storing the entire model file in a cloud environment, increasing the risk of data leakage and computing power theft. To improve security, model shard files can be deployed in local data centers, leveraging existing security measures to protect the model and data.

[0071] In some embodiments, since shard files consume less storage resources, at least a portion of the shard files can be deployed locally in a medical institution or a financial data center, thereby utilizing existing security measures to prevent hackers from stealing computing power and ensure that the user's model-specific services are safe and controllable.

[0072] Specifically, you can choose to deploy shard files across local data centers and cloud environments based on your organization's security needs and resource availability. For example, you could deploy some shard files on a healthcare institution's local servers and others in a financial data center. Ensure the security of each deployment location and protect the shard files using existing firewalls, encryption, and access control policies.

[0073] In some embodiments, multi-factor authentication (MFA) and role-based access control (RBAC) can be used to further enhance security. Shard files can also be stored encrypted to ensure that even if a shard file is illegally accessed, it cannot be directly read or used. For example, a hardware security module (HSM) or encryption key management system (KMS) can be used to manage encryption keys.

[0074] In some embodiments, because financial institutions and medical institutions implement strict real-time monitoring and logging of access to external services, they can promptly detect and respond to abnormal access behavior. They also conduct regular security audits based on regulatory requirements. Therefore, they can integrate existing security systems of financial institutions and medical institutions to check the integrity and security of shard files.

[0075] In some embodiments, different departments of a hospital can share a portion of the medical knowledge database or smart services, and privately deploy shard files based on the characteristics of their own departments, thereby avoiding taking up too much space in the cloud. Based on the computing power of the shard files, they can adaptively localize or professionally upgrade commonly used services such as medical guidance assistance and medical image analysis.

[0076] For example, model shard files for different users can be stored separately in local data centers of medical institutions and financial regulatory agencies, ensuring that each user's service runs independently without interfering with each other. Virtualization or containerization technologies can also be used to further isolate the model operating environment. Performance monitoring based on virtualization or containerization technologies can be implemented to ensure that the loading and inference processes of shard files meet performance requirements. Based on actual needs, the number and size of shard files can be dynamically adjusted, and at least a portion of the shard files and corresponding computing tasks can be promptly migrated to the cloud to optimize model performance and resource utilization.

[0077] In some embodiments, sharding or quantization can be combined to further miniaturize model files, and even enable localized version management of model files and image files. For example, distributed version control tools can be used to manage model file and runtime environment versions, enabling seamless upgrades and rollbacks. By implementing version control for model files and the runtime environment, rapid rollback to previous versions is ensured when necessary.

[0078] S103: Delete the target large model in the first container, and generate an image file based on the first container.

[0079] It should be understood that the target large model can run in the first container because, on the one hand, the first container's own establishment process has equipped the target large model with sufficient hardware resources, which are solidified as parameters of the first container to facilitate container management by the device; and, on the other hand, the first container, through step S101, has established and configured an operating system and dependency tree that are adaptable to the target large model. In particular, because large models themselves have extremely high performance requirements, they require diverse and complex control optimization of computer hardware capabilities. The operating systems widely used in medical institutions such as hospitals, research institutes, and clinics, as well as financial institutions such as banks, securities firms, and insurance companies, lack the ability to directly run large models. In traditional specific deployment processes, they must rely on professional adaptation. The product of this adaptation is a series of software, tools, scripts, and files known as a dependency tree. Therefore, by deleting the target large model in the first container and regenerating the image file, the corresponding dependency tree, operating system and other contents can be retained in the image file. When this device or other devices need to redeploy the target large model, as long as the device can provide enough hardware resources required by the first container, it can directly expand and establish a deployment container through the image file, and the target large model can be "ready to move in". There is no need for professionals to re-adapt and optimize the operating system and dependency tree that can adapt to the target large model based on the basic parameters of the target large model.

[0080] It should be further understood that since the image file is generated based on the first container, the generator can intuitively obtain the hardware resources required by the container based on the container technology, thereby obtaining the hardware configuration parameter ranges such as "available", "smooth", and "preferred". For the deployer and consumer of the model, they can obtain whether their current hardware resource environment can adapt to the model file based on the generator's prompts or container technology, thereby avoiding the problem that the hardware resource environment is not sufficient for smooth operation, but under the premise of not knowing the actual required configuration of the large model and its container, it takes a lot of time to load and deploy the container and model, and it is very likely to be proved to be an unfeasible technical problem. For example, the deployer and consumer of the model can quickly determine whether their current hardware resource environment can adapt to the model file based on the generator's prompts or container technology. This avoids performance problems caused by insufficient hardware resources, especially when the actual required configuration of the large model and its container is unknown, avoiding the problem of taking a lot of time to load and deploy the container and model, which may eventually fail due to insufficient hardware resources.

[0081] In some embodiments, during the process of establishing the first container, the developer can equip the target large model with sufficient hardware resources, and further solidify these resources as parameters of the first container so that devices and users can manage the container. Specifically, these hardware resources include but are not limited to CPU, memory, GPU, etc. The configuration parameters of these resources can be clearly recorded in the container's configuration file to ensure that the container can automatically allocate the corresponding resources when it starts. For example, for a large model that requires high-performance computing, the first container may be configured with multiple GPU cores and a large amount of memory. These configuration parameters are clearly set when the container is created and remain unchanged during subsequent operations.

[0082] In some embodiments, the operating system, the dependency tree, and the application interface can be packaged and compressed to generate an image file. It should be understood that the dependency tree contained in the image file ensures that the large model can run normally in the corresponding operating system environment, including but not limited to a specific version of the interpreter, deep learning framework, and various dependent libraries. The application interface may include an application layer, which may include a prompt word module, a security module, and / or an interactive interface module.

[0083] S104: Associating the model file with the image file to obtain associated information data.

[0084] It should be understood that the associated information data may include model version management data and image version management data. Based on this data, the corresponding relationship between the model file and the image file can be obtained. Furthermore, during storage, the version management data can be stored simultaneously with the model file or image file to ensure bidirectional traceability, prevent tampering and accidental modification, and safeguard the security and reliability of the model file and image file.

[0085] In some embodiments, the associated information data may include first associated information data and second associated information data. When the first container is operated, the associated information data may only include the first associated information data, and after the associated information data is transferred to the cloud storage device, the cloud storage device generates the first associated information data based on a specific algorithm. The first associated data can be written by the user himself, or it can be automatically generated by the software program of the terminal device. Among them, the first associated information data may include data such as numbering marks, file names, and specific associated remarks information, and may also include some basic model version management data and image version management data. The second associated information data may include data such as type index, name index, consistency key, and some advanced model version management data and image version management data.

[0086] It should be understood that version management is a complex task that is combined with manual operations and performed by at least one device. Version management data can also be split into multiple parts, each of which is interdependent and can be generated by different devices. For example, based on the name or container identifier filled in by the user, the first associated data such as the file name and number can be generated on the user side, and then the first associated data can be uploaded to a server such as a cloud storage device. Then, based on the existing storage network environment, mutually pointing hyperlink data can be generated based on the first associated data such as the file name filled in by the user. Second associated data such as a consistency verification code can also be generated based on the model file and the image file itself.

[0087] S105: Transmit the associated information data, model file, and image file to a cloud storage device.

[0088] It should be understood that the cloud storage device and the device performing steps S101-S104 can be the same device or different devices. For example, the cloud storage device can include a processor for performing steps S101-S104. In another example, after performing steps S101-S104 using the cloud storage device, the model file and the image file can be transferred to the storage functional unit of the cloud storage device for storage.

[0089] In some embodiments, when the associated information data including the first associated information data is uploaded to the cloud storage device, the cloud storage device can further establish indexes such as type and name and generate second associated information data such as consistency keys based on the first associated information data. In this way, on the one hand, the queryability and operability of the model file and / or the image file can be improved by establishing indexes, which facilitates ordinary users of relevant institutions in the medical and elderly care health fields to adapt and deploy according to the actual hardware environment, and also facilitates banks, insurance companies, securities companies and other institutions in the financial technology field to adaptively search and select models based on the data structure characteristics of different data sources; on the other hand, the security and verifiability of cloud storage can be further improved by generating consistency keys. It should be understood that the task of establishing indexes such as type and name and generating second associated information data such as consistency keys is handed over to cloud storage devices such as servers for execution, which helps users save computing power loss of local terminals when uploading the image files and model files, and hands over more professional and confidential key generation tools or index establishment and other computing power-consuming tasks to centralized or semi-centralized computing centers for unified processing, thereby improving the fluency and security of cloud services and model management distribution and deployment application systems.

[0090] It should be understood that by separating model files and image files for compression, cloud storage and deployment, it is possible to apply reliable compression methods to each of them. Secondly, the image file (containing the dependency tree and the specific operating system) is associated with the model file for version management, decoupled storage and unified deployment. This not only makes it convenient for deployment personnel to develop and debug large models, saving time for installing and debugging dependency trees and operating systems, but also facilitates development and adaptation on the storage side, and decouples storage to save resources.

[0091] In some embodiments, the cloud storage device may include a first cloud storage system and a second cloud storage system. When transferring associated information data, model files, and image files to the cloud storage device, the model version management data and model files may be transferred to the first cloud storage system, and the image version management data and image files may be transferred to the second cloud storage system. It should be understood that the first cloud storage system can perform storage optimization for the specific data structure of the model file, and the second cloud storage system can perform storage optimization for the specific data structure of the image file. It can also implement associated query, associated download, and associated verification of the model file and the image file based on associated information data such as version management data, thereby solving problems such as mixed data structures and low storage efficiency during hybrid storage, and ensuring the matching and association between the model file and the image file in the subsequent deployment process, greatly facilitating users to independently deploy and install the model, and ensuring accurate and secure matching between the model file and the image file. It should be understood that the matching relationship between the model file and the image file may include one-to-one, many-to-one, one-to-many, many-to-many, etc. relationships, further improving deployment flexibility and reducing the computing power consumption of the deployment optimization algorithm of the automated scaling system. The appropriate image file or model file can be directly selected for deployment based on the corresponding relationship.

[0092] Furthermore, the cloud storage device may include a first cloud storage device and a second cloud storage device, with the first cloud storage system running on the first cloud storage device and the second cloud storage system running on the second cloud storage device. It should be understood that the file capacities of image files and model files are often not consistent. In the field of small models, the file capacity of image files is often much larger than that of model files, while in the field of large models, the file capacity of image files is often much smaller than that of model files. Furthermore, a single image file may be used for multiple model files, resulting in the total file size of image files being smaller than the total file size of model files. Simultaneously releasing image files and model files on the same device will occupy a large amount of bandwidth resources and may even cause malfunctions, data corruption, and version management failure. Therefore, a specific analysis can often be conducted on a case-by-case basis. For example, for specific model architectures in vertical fields such as fintech and healthcare, as well as the deployment environment of cloud service providers, storage space and storage system algorithms can be optimized, and different cloud storage devices can be selected for classified storage to ensure smooth version management and model deployment. For another example, model files for certain large models can be quantized and stored in shards.

[0093] This application also proposes a large model deployment method, which can be configured on a computer device such as a terminal, a server, a server cluster, a distributed computing system, etc. The device can be the same device as the computer device that performs steps S101 to S105, or it can be a cloud storage device, or it can be a different device from the computer device that performs steps S101 to S105. Figure 4 , Figure 4 This is a schematic flow chart of the steps of the large model deployment method provided by an embodiment of the present application. Figure 4 As shown, the large model deployment method may specifically include steps S106 to S108.

[0094] S106: Acquire and establish a second container based on the associated information data.

[0095] It should be understood that a device (second device) different from the computer device (first device) that executes steps S101 to S105 can be used to obtain associated information data through a communication network, and based on the hardware parameter requirements, software environment requirements, container operating parameters, etc. in the associated information data, the specific device for establishing the container can be determined to establish a second container.

[0096] In some embodiments, the associated information data can be separated from the image files and model files and stored separately on an intranet node or an extranet node. This ensures smooth and convenient user access to and retrieval of the associated information data. After obtaining the associated information data, operations such as pulling the image files associated with the associated information data and establishing a second container can be performed as needed.

[0097] In some embodiments, you can first establish a second container and then pull and load the image file, or you can first pull and load the image file and then establish the second container. In containerization technology, container establishment and image loading are two key steps. Traditional containerization processes usually require pulling the image file first and then establishing a container based on the image file. However, in environments where rapid deployment is required or in resource-constrained environments, if the image file is large, the pulling process may take a lot of time, affecting deployment efficiency. Therefore, you can also first establish a basic container for basic testing, and then load different image files at the same time or thereafter according to specific needs.

[0098] S106a: Obtain resource demand information.

[0099] It's important to understand that during the deployment of large AI models, the model's resource requirements may change dynamically as the usage scenario evolves. For example, in high-concurrency inference scenarios, more computing resources are needed to meet real-time response requirements; however, in low-load scenarios, excessive resources can lead to waste. Traditional deployment methods often struggle to flexibly adapt to these dynamic changes, resulting in inefficient resource utilization and difficulty meeting real-time requirements.

[0100] In some embodiments, the operating status of the model can be monitored in real time, including but not limited to CPU usage, memory usage, GPU usage, network bandwidth, performance indicators, cluster failure rate, etc., and then the current resource requirements of the model can be analyzed based on the monitoring data, and possible changes in resource requirements in the future can be predicted to obtain resource requirement information. Specifically, prediction and computing resource planning can be performed through machine learning algorithms or preset rule engines. For example, if it is detected that the CPU usage occupied by the model continues to exceed 80%, it is predicted that the model may require more computing resources.

[0101] S106b. Obtain deployment permissions for multiple devices based on resource demand information.

[0102] It should be understood that the amount of resources that need to be increased or decreased can be determined based on resource demand information, and specific equipment can be reduced, selected, or added to meet these needs. For example, by interacting with the management interface of the cloud service provider or local data center, deployment permissions for more equipment can be requested or released. For example, if additional resources are needed, more virtual machine instances can be requested from the cloud service provider; if resources need to be reduced, currently underutilized instances can be released. Different virtual machine instances can also be requested from the cloud service provider, such as concentrating 30% of the original tasks on server cluster A. To save computing power costs, this portion of the tasks can be migrated to server cluster B when it is predicted that computing power costs will decrease over a certain period of time.

[0103] S106c: Establish a second container on multiple devices.

[0104] It should be understood that second containers can be created on multiple devices that have been granted deployment permissions. Since these containers can be quickly started based on pre-built image files, the model can be run quickly on new devices, improving the convenience and efficiency of load balancing.

[0105] In some embodiments, container orchestration tools can be used to quickly deploy and manage container instances on multiple devices to achieve dynamic expansion of the model. The network connection and storage resources of the container can also be further configured to ensure that the model can access necessary data and external services. With this setting, on the one hand, by dynamically adjusting resources, the resource requirements of the model under different loads are met, thus avoiding waste of resources and improving resource utilization efficiency; on the other hand, in high-concurrency scenarios, resources can be quickly expanded to ensure that the model can respond to user requests in a timely manner and meet real-time requirements; on the other hand, by dynamically adjusting resources in an automated manner, the flexibility of deployment is improved, and it can better cope with dynamically changing needs. In addition, automated resource management reduces manual intervention, reduces operation and maintenance costs, improves the stability and reliability of the system, and better meets the data security and information security needs of financial regulatory agencies and sensitive medical data.

[0106] S107: Pull and deploy the image file to the second container through the cloud storage device according to the associated information data.

[0107] It should be understood that because the model image files are stored in cloud storage devices, traditional image file deployment methods often require manual screening operations and lack management of the relationship between image files and containers, and between image files and model files. This not only increases the complexity of deployment but also may lead to deployment errors and waste of resources. The cloud storage device can be a public cloud service or private cloud storage, which is not specifically limited here.

[0108] In some embodiments, to facilitate rapid deployment in different devices and environments, automated tools can be used to pull the appropriate image file from the cloud storage device based on the mapping relationship between the image file and the container in the associated information data, such as the image file's storage location, version information, and dependencies, thereby reducing the error rate of manual operations. Especially for ultra-large computing power clusters, the computing power resources corresponding to different devices are not the same. Traditional deployment methods and migration processes may lead to resource waste or even system failures, such as repeated downloading of image files and underutilized container instances. The use of automated tools can improve matching accuracy and resource utilization efficiency, ensuring the reliability, flexibility, and security of model deployment.

[0109] S108: Pull and deploy the model file to the second container through the cloud storage device according to the associated information data.

[0110] It should be understood that the associated information data may also include associations or conditional requirements between model files and image files, or between model files and containers. This associated information data can be used to analyze and balance these relationships, thereby simplifying the large model deployment process, reducing deployment errors, improving resource utilization efficiency, and enhancing system scalability. In some embodiments, further analysis and balance can be performed based on current conditions such as computing power shortages and user access volume to ensure that the large model deployment process is accurate, efficient, controllable, and flexible.

[0111] See also Figure 5 , Figure 5 1 is a schematic diagram of the structure of a large model deployment device provided in an embodiment of the present application, wherein the large model deployment device is used to execute the aforementioned large model deployment method. The large model deployment device can be configured in a terminal or a server.

[0112] like Figure 5 As shown, the large model deployment device 100 includes a container testing module 101, a model packaging module 102, an image production module 103, a file association module 104 and a storage and transmission module 105.

[0113] The container testing module 101 is used to establish a first container and run the target large model through the first container.

[0114] The model packaging module 102 is used to package the target large model to obtain a model file.

[0115] The image creation module 103 is configured to delete the target large model in the first container and generate an image file based on the first container.

[0116] The file association module 104 is used to associate the model file with the image file to obtain association information data.

[0117] The storage transmission module 105 is used to transmit the associated information data, the model file and the image file to a cloud storage device.

[0118] In some embodiments, the large model deployment apparatus 100 may further include a container deployment module, which may be configured on a different device from the aforementioned modules 101-105. For example, the container testing module 101, model packaging module 102, image creation module 103, file association module 104, and storage transmission module 105 may be configured on a first device, while the container deployment module may be configured on a second device. The container deployment module may be configured to obtain and establish a second container based on the associated information data.

[0119] In some embodiments, the large model deployment device 100 may further include an image deployment module, which may be configured on a device different from the aforementioned modules 101-105. For example, the container testing module 101, the model encapsulation module 102, the image creation module 103, the file association module 104, and the storage transmission module 105 may be configured on a first device, while the image deployment module may be configured on a second device. The image deployment module may be used to obtain and establish a second container based on associated information data. The image deployment module may be used to pull and deploy the image file to the second container through a cloud storage device based on the associated information data.

[0120] In some embodiments, the large model deployment apparatus 100 may further include a model deployment module, which may be configured on a device different from the aforementioned modules 101-105. For example, the container testing module 101, the model encapsulation module 102, the image creation module 103, the file association module 104, and the storage and transmission module 105 may be configured on a first device, while the model deployment module may be configured on a second device. The model deployment module may be used to obtain and establish a second container based on the associated information data. The model deployment module may be used to pull and deploy the model file to the second container via a cloud storage device based on the associated information data.

[0121] It should be noted that those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices and modules can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0122] The above-mentioned device can be realized in the form of a computer program. The computer program can be used in Figure 6 Runs on the computer device shown.

[0123] See also Figure 6 , Figure 6 This is a schematic block diagram of the structure of a computer device provided in an embodiment of the present application. The computer device may be a server. Figure 6 The computer device includes a processor, a memory, and a network interface connected through a system bus, wherein the memory may include a non-volatile storage medium and an internal memory.

[0124] The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions, which, when executed, can cause the processor to execute any of the large model deployment methods of the present application.

[0125] The processor is used to provide computing and control capabilities and support the operation of the entire computer equipment.

[0126] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor, the processor can execute any large model deployment method.

[0127] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art will understand that Figure 6 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0128] It should be understood that the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0129] In some embodiments, the processor is configured to execute a computer program stored in the memory to implement the following steps:

[0130] Establish a first container and run the target large model through the first container; encapsulate the target large model in a grouping manner and / or a quantization manner to obtain a model file in a binary format; delete the target large model in the first container and generate an image file based on the first container; associate the model file with the image file to obtain associated information data, wherein the associated information data includes model version management data and image version management data, and the correspondence between the model file and the image file can be obtained based on the model version management data and the image version management data; transfer the associated information data, the model file and the image file to a cloud storage device.

[0131] In some embodiments, when the processor is used to establish a first container and run the target large model through the first container, it can specifically be used to implement: establishing the first container and running an operating system through the first container; obtaining a target large model that is adapted to the current situation based on the current situation of the first container and its operating system; setting a dependency tree and an application interface through the operating system based on the target large model, wherein the target large model can run on the operating system through the dependency tree, and the target large model can interact through the application interface. Running the target large model through the first container.

[0132] In some embodiments, when the processor is used to implement the generation of the image file according to the first container, it can be specifically used to implement: packaging and compressing the operating system, the dependency tree and the application interface to generate the image file

[0133] In some embodiments, the processor can also be used to implement: obtaining and establishing a second container based on the associated information data; pulling and deploying the image file to the second container through the cloud storage device based on the associated information data; pulling and deploying the model file to the second container through the cloud storage device based on the associated information data.

[0134] In some embodiments, when the processor is used to implement the establishment of the second container, it can specifically be used to implement: obtaining resource requirement information; obtaining deployment permissions for multiple devices based on the resource requirement information; and establishing the second container on the multiple devices.

[0135] In some embodiments, the cloud storage device includes a first cloud storage system and a second cloud storage system.

[0136] In some embodiments, the cloud storage device includes a first cloud storage device and a second cloud storage device, the first cloud storage system runs on the first cloud storage device, and the second cloud storage system runs on the second cloud storage device.

[0137] In some embodiments, when the processor is used to implement the transfer of the associated information data, the model file and the mirror file to the cloud storage device, it can specifically be used to implement: transferring the model version management data and the model file to the first cloud storage system; transferring the mirror version management data and the mirror file to the second cloud storage system.

[0138] A computer-readable storage medium is also provided in an embodiment of the present application. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. The processor executes the program instructions to implement any large model deployment method provided in the embodiment of the present application.

[0139] The computer-readable storage medium may be an internal storage unit of the computer device described in the aforementioned embodiment, such as a hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a SmartMedia Card (SMC), a Secure Digital (SD) card, a flash memory card, etc., equipped on the computer device.

[0140] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present application, and such modifications or substitutions should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A large model deployment method, characterized in that: include: Establishing a first container, and running the target large model through the first container; Encapsulating the target large model in a grouping manner and / or a quantized manner to obtain a model file in a binary format; Deleting the target large model in the first container and generating an image file based on the first container; Associating the model file with the image file to obtain associated information data, wherein the associated information data includes model version management data and image version management data, and a corresponding relationship between the model file and the image file can be obtained based on the model version management data and the image version management data; The associated information data, the model file, and the image file are transmitted to a cloud storage device.

2. The method according to claim 1, characterized in that The step of establishing a first container and running the target large model through the first container includes: Establishing a first container and running an operating system through the first container; According to the current situation of the first container and its operating system, obtaining a target macro model adapted to the current situation; According to the target large model, a dependency tree and an application interface are set through the operating system, wherein the target large model can run on the operating system through the dependency tree, and the target large model can interact through the application interface. The target large model is run through the first container.

3. The method according to claim 2, characterized in that Generating an image file according to the first container includes: The operating system, the dependency tree, and the application interface are packaged and compressed to generate an image file.

4. The method according to claim 1, wherein include: Acquire and establish a second container based on the associated information data; Pulling and deploying the image file to the second container through the cloud storage device according to the associated information data; According to the associated information data, the model file is pulled and deployed to the second container through the cloud storage device.

5. The method according to claim 4, characterized in that The establishing of the second container includes: Obtain resource demand information; Obtaining deployment permissions for multiple devices based on the resource requirement information; A second container is established on the plurality of devices.

6. The method according to claim 1, characterized in that The cloud storage device includes a first cloud storage system and a second cloud storage system; and the transferring of the associated information data, the model file, and the image file to the cloud storage device includes: Transmitting the model version management data and the model file to a first cloud storage system; The image version management data and the image file are transmitted to a second cloud storage system.

7. The method according to claim 6, characterized in that The cloud storage device includes a first cloud storage device and a second cloud storage device. The first cloud storage system runs on the first cloud storage device, and the second cloud storage system runs on the second cloud storage device.

8. A large model deployment device, characterized in that: include: A container testing module, configured to establish a first container and run a target large model through the first container; A model packaging module is used to package the target large model to obtain a model file; an image creation module, configured to delete the target macro model of the first container and generate an image file based on the first container; A file association module, associating the model file with the image file to obtain association information data; The storage transmission module transmits the associated information data, the model file and the image file to a cloud storage device.

9. A computer device, characterized in that: The computer device includes a processor, a memory, and a computer program stored in the memory and executable by the processor, wherein when the computer program is executed by the processor, the large model deployment method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, enables the processor to implement the large model deployment method according to any one of claims 1 to 7.