Mirror image management method and device, electronic equipment and storage medium
By building the basic business image layer and dynamic dependency component resource library, the basic image and dependency components are decoupled, and the dependency components are loaded on demand and image real-time updates are solved, and the problems of bloated images, low update efficiency and frequent version conflicts in large-scale training image management are achieved, and efficient storage and operation and maintenance are achieved.
Patent Information
- Application Number
- CN202510236374.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-05-30
AI Technical Summary
In the prior art, large-model training mirror management has problems such as bloated image, inefficient update efficiency and frequent version conflicts, resulting in swelling storage resources, high operation and maintenance costs and slow platform response speed.
By building the basic business image layer and dynamic dependency component resource library, the basic image and dependency components are decoupled, and the dependency components are loaded on demand and updated in real time in mirroring, reducing the image storage overhead and operation and maintenance complexity.
It realizes order of magnitude optimization of storage overhead, reduces operation and maintenance costs, improves platform response speed and iteration efficiency, and ensures maximum resource utilization and overall system performance.
Smart Images

Figure CN120066532A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of cloud computing technology, and in particular, to a mirror management method, apparatus, electronic device, and storage medium. Background Art
[0002] In related technologies, AI (Artificial Intelligence) platforms generally adopt a containerized deployment solution, that is, by constructing independent container images to meet the dependency requirements of each large model for a customized training environment. The direct problem brought about by this solution is the exponential expansion of storage resources. More seriously, when basic components need to be upgraded, maintenance personnel need to update all relevant images one by one, and the operation complexity and error probability increase linearly with the number of images. In the field of environment isolation technology, the Conda (open-source package management and environment isolation tool) virtual environment solution has been tried to optimize this problem. By creating multiple Conda environments within a single basic image, it is theoretically possible to achieve dependency isolation for different models. However, in practical applications, it is found that this solution causes the volume of the basic image to expand to more than 50 GB, and there are the following defects:
[0003] 1. Low mirror transmission efficiency: When scheduling model training tasks, even if only a small part of the dependencies in the environment are used, the entire giant mirror still needs to be fully loaded.
[0004] 2. Risk of version conflicts: Multiple Conda environments share the underlying system libraries. When the system library versions required by different environments are incompatible, runtime errors that are difficult to troubleshoot may be triggered.
[0005] 3. Difficulty in security updates: Security patches in the basic image need to rebuild all Conda environments, and the update process may break the dependency relationships of existing environments.
[0006] 4. Severe storage redundancy: The common dependencies between different Conda environments are stored repeatedly in the image, resulting in waste of storage space. Summary of the Invention
[0007] This application provides a mirror management method, apparatus, electronic device, and storage medium to at least solve the problem of high complexity in mirror management in related technologies.
[0008] This application provides a mirror management method, including:
[0009] Construct a basic business mirror layer, which is used to provide basic components and general environment configurations to support large model training.
[0010] Obtain a target large model and at least one dependent component corresponding to the target large model, and generate a one-to-one mapping relationship based on the target large model and the dependent component corresponding to the target large model.
[0011] Determine the installation package of the dependent component corresponding to the target large model based on the one-to-one mapping relationship, and generate a dependent component resource library based on multiple installation packages of the dependent components;
[0012] In response to receiving a large model training request, parse the large model training request, and based on the parsing result, extract the target dependent component installation package from the dependent component resource library;
[0013] Install the target dependent component installation package on the basic service image layer to generate a target service image for training the large model, so as to respond to the large model training request.
[0014] This application also provides an image management device, including:
[0015] A basic service image layer construction module, configured to construct a basic service image layer, and the basic service image layer is used to provide basic components and general environment configurations that support large model training;
[0016] A large model configuration module, configured to obtain a target large model and at least one dependent component corresponding to the target large model, and generate a one-to-one mapping relationship based on the target large model and the dependent component corresponding to the target large model;
[0017] A dependent component resource library construction module, configured to determine the installation package of the dependent component corresponding to the target large model based on the one-to-one mapping relationship, and generate a dependent component resource library based on multiple installation packages of the dependent components;
[0018] A request parsing module, configured to, in response to receiving a large model training request, parse the large model training request, and based on the parsing result, extract the target dependent component installation package from the dependent component resource library;
[0019] A request response module, configured to install the target dependent component installation package on the basic service image layer to generate a target service image for training the large model, so as to respond to the large model training request.
[0020] This application also provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of the image management method when executing the computer program, including:
[0021] Construct a basic service image layer, and the basic service image layer is used to provide basic components and general environment configurations that support large model training;
[0022] Obtain a target large model and at least one dependent component corresponding to the target large model, and generate a one-to-one mapping relationship based on the target large model and the dependent component corresponding to the target large model;
[0023] Based on a one-to-one mapping relationship, determine the dependency component installation packages corresponding to the target large model, and generate a dependency component resource library based on multiple dependency component installation packages;
[0024] In response to receiving a large model training request, parse the large model training request, and based on the parsing result, extract the target dependency component installation package from the dependency component resource library;
[0025] Install the target dependency component installation package on the basic service image layer to generate a target service image for training the large model, so as to respond to the large model training request.
[0026] This application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of the image management method are implemented, including:
[0027] Build a basic service image layer, which is used to provide basic components and general environment configurations that support large model training;
[0028] Obtain the target large model and at least one dependency component corresponding to the target large model, and generate a one-to-one mapping relationship based on the target large model and the dependency component corresponding to the target large model;
[0029] Based on the one-to-one mapping relationship, determine the dependency component installation packages corresponding to the target large model, and generate a dependency component resource library based on multiple dependency component installation packages;
[0030] In response to receiving a large model training request, parse the large model training request, and based on the parsing result, extract the target dependency component installation package from the dependency component resource library;
[0031] Install the target dependency component installation package on the basic service image layer to generate a target service image for training the large model, so as to respond to the large model training request.
[0032] This application also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the image management method are implemented, including:
[0033] Build a basic service image layer, which is used to provide basic components and general environment configurations that support large model training;
[0034] Obtain the target large model and at least one dependency component corresponding to the target large model, and generate a one-to-one mapping relationship based on the target large model and the dependency component corresponding to the target large model;
[0035] Based on the one-to-one mapping relationship, determine the dependency component installation packages corresponding to the target large model, and generate a dependency component resource library based on multiple dependency component installation packages;
[0036] In response to receiving a large model training request, parse the large model training request, and based on the parsing result, extract the target dependent component installation package from the dependent component resource library;
[0037] Install the target dependent component installation package on the basic service image layer to generate a target service image for training the large model, so as to respond to the large model training request.
[0038] Through the image management method, device, electronic device and storage medium provided by this application, decouple the basic service image and dynamic dependent components, enabling the storage overhead to be optimized by an order of magnitude, eliminating the duplicate storage of common dependencies, reducing the operation and maintenance costs. Through the mechanism of pre-setting the basic image and loading dependent components on demand, compress the startup delay time, reduce the data transfer volume of a single environment update operation, improve the platform response speed and iteration efficiency. The general design of the basic service image ensures the maximization of resource utilization, avoids resource waste caused by environmental differences, and loading dependent components on demand further reduces unnecessary storage space occupation and improves the overall system performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0040] Figure 1 It is an application environment diagram of an image management method provided by an embodiment of the present application;
[0041] Figure 2 It is a flowchart of an image management method provided by an embodiment of the present application;
[0042] Figure 3 It is a schematic diagram of module integration of an image management method provided by an embodiment of the present application;
[0043] Figure 4 It is a flowchart of constructing a basic service image of an image management method provided by an embodiment of the present application;
[0044] Figure 5 It is a flowchart of the working process of a dynamic expansion construction engine of an image management method provided by an embodiment of the present application;
[0045] Figure 6 It is a structural block diagram of an image management device provided by an embodiment of the present application;
[0046] Figure 7 It is an internal structure diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0047] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts fall within the protection scope of the present application.
[0048] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0049] It should be noted that the terms "S1", "S2", etc. are only used for the purpose of describing steps and do not particularly refer to the meaning of order or sequence. Nor are they used to limit the present application. They are only used to facilitate the description of the method of the present application and cannot be understood as indicating the order of steps. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present application.
[0050] With the rapid development of artificial intelligence technology, large-scale pre-trained models have become the core driving force for promoting industry transformation. The field of large models is experiencing unprecedented rapid development, with the heat continuing to rise. Many enterprises, institutions, organizations, and universities have thrown themselves into the research, development, and application of large models, regarding them as the key force to promote their own business innovation, enhance competitiveness, and expand the boundaries of scientific research. Based on this, there is a need for the architecture innovation of enterprise-level AI platforms, that is, the platform needs to support the parallel training, evaluation, and deployment operations of multiple large models simultaneously, forming a complex technical integration system.
[0051] In a typical large model production process, the construction of a training environment is a fundamental technical challenge, and there are significant technical differences in the large models developed by various institutions: at the software level, different models have different requirements for the versions of deep learning frameworks such as PyTorch (an open-source deep learning framework). This fragmentation of the technical ecosystem leads to the need for a customized training environment for each large model. The relevant technology meets the dependency requirements of each model by constructing independent container images. Among them, mainstream AI platforms generally adopt a containerized deployment solution. Taking the Kubernetes (a cluster management system for containerized applications) cluster management as an example, when N large models need to be supported, the platform needs to maintain N independent images, and each image contains a complete operating system, basic libraries, and model-specific dependencies. The direct problem brought by this architecture is the exponential expansion of storage resources, that is, the volume of a single image usually reaches 15 - 30GB, and supporting 50 models will generate nearly 1.5TB of image storage overhead. More seriously, when basic components (such as NVIDIA drivers) need to be upgraded, maintenance personnel need to update all relevant images one by one, and the operation complexity and error probability increase linearly with the number of images; According to the background technology, in the field of environment isolation technology, the Conda virtual environment solution has been tried to optimize this problem. By creating multiple Conda environments within a single base image, it is theoretically possible to achieve dependency isolation for different models. For example, a base image containing CUDA 12.1 is built, and two environments, conda_env_A (containing PyTorch 2.2 + Python 3.10) and conda_env_B (containing PyTorch 2.4 + Python 3.11), are established respectively. However, in practical applications, it is found that this solution causes the volume of the base image to expand to more than 50GB and has various defects.
[0052] To solve the technical problems such as bloated images, low update efficiency, and frequent version conflicts in the above-mentioned large model training image management, this application provides an image management method, device, electronic device, and storage medium. By decoupling the base service image and dynamic dependency components, the storage overhead is optimized by an order of magnitude, the repeated storage of common dependencies is eliminated, and the operation and maintenance costs are reduced. Through the mechanism of pre-setting the base image and loading dependency components on demand, the startup delay time is compressed, the data transfer volume of a single environment update operation is reduced, and the platform response speed and iteration efficiency are improved. The general design of the base service image ensures the maximum utilization of resources and avoids resource waste caused by environmental differences. Loading dependency components on demand realizes the precise scheduling and efficient reuse of multi-version dependency components, further reducing unnecessary storage space occupation and improving the overall performance of the system.
[0053] To enable those skilled in the art of this technology to better understand the solution of this application, the following further elaborates on this application in combination with the accompanying drawings and specific implementation manners.
[0054] The mirror management method provided by this application can be applied to an application environment as shown in Figure 1 the following. Among them, the terminal 102 communicates with the data processing platform set on the server 104 through the network. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices, and the server 104 can be implemented by an independent server or a server cluster composed of multiple servers.
[0055] An embodiment of this application provides a mirror management method as shown in Figure 2 the following. Taking the terminal in Figure 1 as an example, the method includes the following steps:
[0056] S1: Construct a basic service mirror layer, which is used to provide basic components and general environment configurations for supporting large model training.
[0057] It should be noted that the basic service mirror layer refers to the bottommost basic mirror layer in the Docker mirror. As the base for large model training, the basic service mirror contains all the necessary basic components and general environment configurations for large model training. These basic components and general environment configurations include, but are not limited to, operating systems, Python interpreters, CUDA / cuDNN libraries, common data processing libraries (such as Pandas, NumPy), etc. The basic service mirror layer, as a static basic layer, is an immutable infrastructure for the training environment.
[0058] In some specific implementation manners, as shown in Figure 4 the following, constructing the basic service mirror layer includes:
[0059] Select a target operating system, load the kernel and integrate hardware drivers;
[0060] Based on the static compilation mechanism, pre-install cross-model general components to generate the basic service mirror layer. Among them, the cross-model general components include deep learning acceleration libraries, mathematical calculation libraries, and distributed training frameworks, etc. Static compilation refers to the process of compiling source code into machine code (i.e., executable code) and saving it as a binary file before the program is executed.
[0061] Specifically, a layered container image building technology is used to build the basic business image, which includes an operating system kernel layer and a common dependency library layer. The operating system kernel layer generates a minimized kernel by selecting the operating system required for large model training, trimming unnecessary components, integrating hardware drivers such as NVIDIA GPU drivers, building the operating system based on a customized Linux distribution, and integrating hardware driver modules and security reinforcement components. The common dependency library layer pre-poses components such as deep learning acceleration libraries, mathematical calculation libraries, and distributed training frameworks that are common across models through static compilation, and ensures compatibility. Based on the operating system kernel layer and the common dependency library layer, a corresponding basic business image in the Docker container format is output, and a static basic layer, that is, the basic business image layer, is generated.
[0062] In the above embodiment, by building the basic business image layer, it is ensured that all supported large models can run in a stable and consistent environment, avoiding training failures or performance degradation caused by environmental differences.
[0063] S2: Obtain the target large model and at least one dependency component corresponding to the target large model, and generate a one-to-one mapping relationship based on the target large model and the dependency component corresponding to the target large model.
[0064] It should be noted that a large model refers to a deep learning model with a parameter scale exceeding tens of billions and trained based on a large amount of data, having natural language understanding, generation, and complex reasoning capabilities. These models are usually constructed by deep neural networks, and their design purpose is to improve the expression ability and prediction performance of the model so that it can handle more complex tasks and data; the target large model refers to the machine learning model supported by the system; the dependency component refers to the component required for building the large model training environment.
[0065] In some specific embodiments, obtaining the target large model and at least one dependency component corresponding to the target large model, and generating a one-to-one mapping relationship based on the target large model and the dependency component corresponding to the target large model includes:
[0066] Obtain multiple large models supported by the basic business image layer and the unique identifier corresponding to the target large model, where the unique identifier refers to the large model ID identifier used to identify the corresponding large model instance;
[0067] Obtain at least one dependency component corresponding to the target large model and the attribute information of the dependency component, where the attribute information of the dependency component includes detailed information such as component name and version number;
[0068] Based on the attribute information of the dependency component, generate the dependency item corresponding to the dependency component, that is, integrate information such as component name, version constraint, and installation mode into the dependency item corresponding to the dependency component;
[0069] Map the unique identifier corresponding to the target large model and the dependencies corresponding to the dependent components one by one to generate a one-to-one mapping relationship between the target large model and the dependent components corresponding to the target large model;
[0070] After generating the one-to-one mapping relationship, the method further includes:
[0071] Obtain the programming language version, deep learning framework version, and distributed training tool version corresponding to the target large model, that is, determine the runtime environment declaration by specifying the programming language version, deep learning framework version, and distributed training tool version;
[0072] Generate a large model configuration table based on multiple one-to-one mapping relationships, as well as the programming language version, deep learning framework version, and distributed training tool version corresponding to the target large model.
[0073] Specifically, the large model configuration table at least includes a model ID identifier, a runtime environment declaration, and a dependency list generated based on the dependencies. Use a declarative configuration file in the YAML (a readable data serialization format) format to build an accurate mapping relationship between the large model and the dependent components. The YAML configuration file contains the following core fields:
[0074] model_profile:
[0075] model_id:qwen-72b
[0076] runtime:
[0077] python:3.10.12
[0078] pytorch:2.2.2
[0079] dependencies:
[0080] -name:transformers
[0081] version:4.41.2
[0082] install_mode:pip
[0083] -name:flash-attn
[0084] version:2.5.8
[0085] install_mode:prebuilt
[0086] Among them, the installation mode (install_mode) supports pip, prebuilt (pre-compiled package), and source (source code compilation).
[0087] In the above embodiment, by setting up the large model configuration table to record the mapping relationship between each supported large model and its required dependent components, when the model dependencies change, it is not necessary to rebuild the entire image, and only the corresponding entries of the configuration information need to be updated. In this way, the rapid configuration and flexible adjustment of the training environment are achieved, the preparation work of the training environment is simplified, and the maintenance cost caused by the change of dependent components is effectively reduced.
[0088] S3: Based on the one-to-one mapping relationship, determine the dependent component installation packages corresponding to the target large model, and generate a dependent component resource library based on the multiple dependent component installation packages.
[0089] It should be noted that the dependent component resource library includes multiple dependent component installation packages, and each dependent component installation package is compressed and stored in an independent tar.gz format and is accompanied by its corresponding metadata file.
[0090] In some specific embodiments, generating a dependent component resource library based on multiple dependent component installation packages includes:
[0091] Standardize the dependent component installation packages. Among them, the standardized dependent component installation packages are accompanied by the metadata files corresponding to the dependent components. The metadata files include compatibility identifiers, hash values, etc. The file contains the following core fields:
[0092] {
[0093] "checksum":"sha256:9a8b7c6d...",
[0094] "abi_compatibility":["numpy>=1.22.4,<2.0.0","pandas>=2.0.0"]
[0095] }
[0096] Generate a dependent component resource library based on the standardized multiple dependent component installation packages.
[0097] In the above embodiment, by constructing a dependent component resource library, the dependent component resource library contains the installation packages of different versions of all dependent components required for large model training. These installation packages are pre-compiled and tested to ensure compatibility with the basic business image and the training framework project. When preprocessing the training environment according to the large model configuration module, the installation packages in the dependent component resource library will be extracted and installed as needed, thus avoiding unnecessary resource waste and time consumption.
[0098] S4: In response to receiving a large model training request, parse the large model training request, and based on the parsing result, extract the target dependent component installation package from the dependent component resource library.
[0099] It should be noted that the large model training request includes a model ID, which is set by the user when uploading the large model training task. The parsing result of the large model training request includes the configurations required for the large model training task, such as the large model ID and its corresponding dependent components.
[0100] In some specific embodiments, extracting the target dependent component installation package from the dependent component resource library based on the parsing result includes:
[0101] Obtain the parsing result of the large model training request, where the parsing result at least includes the large model unique identifier;
[0102] Based on the large model configuration table, determine the dependencies corresponding to the large model unique identifier;
[0103] Based on the dependent component resource library, match and extract the target dependent component installation package corresponding to the dependencies.
[0104] In some specific embodiments, after extracting the target dependent component installation package from the dependent component resource library, it further includes:
[0105] Obtain the metadata file corresponding to the target dependent component installation package, where the metadata file at least includes a compatibility identifier and a hash value;
[0106] Based on the compatibility identifier and the hash value, verify the target dependent component installation package.
[0107] Specifically, as Figure 5 shown, the large model training request is parsed by the constructed dynamic extension construction engine, and the dynamic extension construction engine executes the following process when the training task is started: According to the model ID specified when the task is submitted, extract the dependencies from the configuration file; initiate a batch query request to the dependent component resource library to obtain the dependent component installation package corresponding to the dependencies; retrieve the dependent component installation package from the resource library to obtain the metadata file of the required dependent component to verify the checksum (integrity check, based on the hash value check) and perform a compatibility check according to the compatibility identifier; install the dependent component packages in sequence on the basis business image to build the large model training environment; output the target business image that meets the large model training.
[0108] In the above embodiments, the constructed dynamic expansion construction engine is responsible for dynamically constructing an isolated training environment based on the static basic layer when the training task of the large model is started, and realizing the construction of the dynamic expansion layer through this engine, so as to realize the "static basic layer + dynamic expansion layer" architecture of the large model training environment. By loading dependent components on demand and isolating the environment, the efficiency of image management is improved, and the multi-model support ability and resource utilization efficiency of the AI platform are enhanced.
[0109] In some specific embodiments, the above method further includes:
[0110] Select a target hash function, where the hash function can include SHA-256, MD5, etc., and can be selected according to actual needs;
[0111] Use the selected hash function to calculate the hash value of the original data corresponding to the required dependent components, and generate a hash value of a fixed length;
[0112] Map the original data and the hash value and store the generated mapping relationship;
[0113] When integrity verification is required, extract the original data corresponding to the required dependent components, and calculate the hash value of the original data through the above-selected target hash function to obtain the current hash value calculation result;
[0114] Compare the current hash value calculation result with the received hash value to obtain a comparison result;
[0115] In response to the comparison result being inconsistent, it is determined that the integrity of the data is low, re-obtain the required dependent components, and repeat the integrity verification until the integrity standard is met;
[0116] In response to the comparison result being consistent, it is determined that the integrity of the data is high. At this time, perform compatibility verification according to the compatibility identifier to determine the final dependent component installation package.
[0117] In the above embodiments, the integrity verification of the data corresponding to the dependent component installation package is performed through a hash function, which can improve the accuracy of the generation result of the target service image used for training the large model later.
[0118] S5: Install the target dependent component installation package on the basic service image layer to generate a target service image for training the large model in response to the large model training request.
[0119] It should be noted that the target service image refers to the large model training environment image. As described in the above steps, the task of installing the target dependent component installation package on the basic service image layer is parsed and executed by the dynamic expansion construction engine.
[0120] In some specific embodiments, the above method further includes:
[0121] Obtaining the core code logic of the training framework corresponding to the target large model;
[0122] Generating a training framework project file based on the core code logic to construct a training framework project layer.
[0123] Specifically, as Figure 3 shown, the training framework project layer supports multiple programming languages and frameworks to adapt to the preferences and needs of different developers. It is independently managed in the form of a code repository to achieve physical isolation between the framework code and the base image. The training framework project module carries the core code logic of large model training. The flexibility of this module is crucial because with the rapid development of AI technology, new training algorithms and optimization strategies are constantly emerging and need to be updated frequently to adapt to the latest research results. By dynamically loading the training framework project as an independent module, it allows for rapid iteration of the framework code without affecting other components, thus accelerating the application and implementation of new algorithms.
[0124] In the above embodiment, by constructing the target business image and selecting the corresponding training framework project file to run the target business image to execute the large model training task, creating a training framework project module containing the core code logic of large model training, supporting independent updates of the framework code to adapt to the latest training algorithms and optimization strategies, improving the flexibility and scalability of image management, and ensuring efficient update and maintenance.
[0125] In the above mirror management method, the method includes: constructing a basic service mirror layer, which is used to provide basic components and general environment configurations for supporting large model training; obtaining a target large model and at least one dependent component corresponding to the target large model, and generating a one-to-one mapping relationship based on the target large model and the dependent components corresponding to the target large model; determining an installation package of the dependent component corresponding to the target large model based on the one-to-one mapping relationship, and generating a dependent component resource library based on multiple installation packages of the dependent components; in response to receiving a large model training request, parsing the large model training request, and extracting a target dependent component installation package from the dependent component resource library based on the parsing result; installing the target dependent component installation package on the basic service mirror layer to generate a target service mirror for training the large model, so as to respond to the large model training request. The beneficial effects of this application include: (1) Efficient update and maintenance: Through modular design, this application decouples the static components and dynamic dependent components for constructing the large model training environment mirror, and decomposes them into multiple independent and configurable components. When the model dependencies are updated or the framework is upgraded, only the changed parts need to be accurately updated, avoiding the reconstruction of the overall mirror and the cumbersome large file copying process, improving the update efficiency, and also enabling the storage overhead to be optimized by an order of magnitude, eliminating the repeated storage of common dependent components, and reducing the operation and maintenance costs; (2) Optimal utilization of resources: The general design of the basic service mirror ensures the maximization of resource utilization, avoiding resource waste caused by environmental differences. At the same time, the on-demand extraction mechanism of the dependent component resource package further reduces the unnecessary storage space occupation and improves the overall system performance. The startup delay time is compressed through the basic mirror presetting and dependent on-demand loading mechanism. At the same time, the time-consuming for framework update is also compressed. When the version of the model dependent component changes, the operation and maintenance personnel only need to maintain the large model configuration module without triggering the mirror reconstruction process to achieve immediate effect. The on-demand loading mechanism of the dependent components reduces the data transfer volume of a single environment update operation, improving the platform response speed and iteration efficiency; (3) Flexibility and scalability: The introduction of the large model configuration module makes the configuration of the training environment flexible and easy to expand. Whether adding new model support or adjusting the version of the dependent components, it can be achieved by simply updating the configuration table without major adjustments to the mirror structure; (4) By introducing a dynamic extension construction engine, it is ensured that each large model training task uses an independent space, solving the problem of multi-version dependency conflicts. While ensuring environmental isolation, it significantly improves the mirror management efficiency, reduces the mirror storage overhead, compresses the time-consuming for dependent component updates, reduces the startup delay of training tasks, and improves the overall system performance.
[0126] It should be understood that although Figures 2 - 5The steps in the flowchart are shown in sequence according to the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figures 2 - 5 At least a part of the steps in Figures 2 - 5 may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0127] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.
[0128] The embodiments of the present application also provide a mirror management device, as Figure 6 shown, including a basic service mirror layer construction module, a large model configuration module, a dependent component resource library construction module, a request parsing module, and a request response module.
[0129] The basic service mirror layer construction module is used to construct a basic service mirror layer, and the basic service mirror layer is used to provide basic components and general environment configurations to support large model training;
[0130] The large model configuration module is used to obtain a target large model and at least one dependent component corresponding to the target large model, and generate a one-to-one mapping relationship based on the target large model and the dependent components corresponding to the target large model;
[0131] The dependent component resource library construction module is used to determine the dependent component installation packages corresponding to the target large model based on the one-to-one mapping relationship, and generate a dependent component resource library based on multiple dependent component installation packages;
[0132] The request parsing module is used to respond to receiving a large model training request, parse the large model training request, and extract the target dependent component installation package from the dependent component resource library based on the parsing result;
[0133] The request response module is used to install the target dependent component installation package on the basic service mirror layer, generate a target service mirror for training the large model, and respond to the large model training request.
[0134] In this embodiment, the mirror management device further includes a training framework project layer construction module, and the training framework project layer construction module is used for:
[0135] Obtain the core code logic of the training framework corresponding to the target large model;
[0136] Based on the core code logic, generate a training framework project file to construct the training framework project layer.
[0137] In this embodiment, constructing the basic business image layer includes:
[0138] Select the target operating system, load the kernel and integrate the hardware driver;
[0139] Based on the static compilation mechanism, preset cross-model general components to generate the basic business image layer.
[0140] In this embodiment, obtain the target large model and at least one dependent component corresponding to the target large model. Based on the target large model and the dependent components corresponding to the target large model, generating a one-to-one mapping relationship includes:
[0141] Obtain multiple large models supported by the basic business image layer and the unique identifier corresponding to the target large model;
[0142] Obtain at least one dependent component corresponding to the target large model and the attribute information of the dependent component;
[0143] Based on the attribute information of the dependent component, generate the dependencies corresponding to the dependent component;
[0144] Perform a one-to-one mapping on the unique identifier corresponding to the target large model and the dependencies corresponding to the dependent components to generate a one-to-one mapping relationship between the target large model and the dependent components corresponding to the target large model;
[0145] After generating the one-to-one mapping relationship, the method further includes:
[0146] Obtain the programming language version, deep learning framework version, and distributed training tool version corresponding to the target large model;
[0147] Based on multiple one-to-one mapping relationships, as well as the programming language version, deep learning framework version, and distributed training tool version corresponding to the target large model, generate a large model configuration table.
[0148] In this embodiment, based on multiple dependent component installation packages, generating a dependent component repository includes:
[0149] Standardize the dependent component installation packages, where the standardized dependent component installation packages are attached with metadata files corresponding to the dependent components;
[0150] Based on the standardized multiple dependent component installation packages, generate a dependent component repository.
[0151] In this embodiment, based on the parsing result, extracting the target dependent component installation package from the dependent component resource library includes:
[0152] Obtain the parsing result of the large model training request, and the parsing result at least includes the unique identifier of the large model;
[0153] Based on the large model configuration table, determine the dependencies corresponding to the unique identifier of the large model;
[0154] Based on the dependent component resource library, match and extract the target dependent component installation package corresponding to the dependencies.
[0155] In this embodiment, the image management device further includes a verification module, and the verification module is used for:
[0156] Obtain the metadata file corresponding to the target dependent component installation package, and the metadata file at least includes a compatibility identifier and a hash value;
[0157] Based on the compatibility identifier and the hash value, verify the target dependent component installation package.
[0158] For the specific limitations of the image management device, reference can be made to the limitations on the image management method in the above text, which will not be elaborated here. Each module in the above image management device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0159] An embodiment of the present application also provides an electronic device, which can be a server, and its internal structure diagram can be as Figure 7 shown. The electronic device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the electronic device is used to store arbitration control data. The network interface of the electronic device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it realizes an image management method.
[0160] Those skilled in the art can understand that Figure 7 the structure shown in is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0161] An embodiment of the present application also provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in the embodiment of the mirror management method, including:
[0162] S1: Construct a basic service image layer, which is used to provide basic components and general environment configurations for supporting large model training;
[0163] S2: Obtain a target large model and at least one dependent component corresponding to the target large model, and generate a one-to-one mapping relationship based on the target large model and the dependent components corresponding to the target large model;
[0164] S3: Based on the one-to-one mapping relationship, determine the installation package of the dependent component corresponding to the target large model, and generate a dependent component resource library based on multiple dependent component installation packages;
[0165] S4: In response to receiving a large model training request, parse the large model training request, and based on the parsing result, extract the target dependent component installation package from the dependent component resource library;
[0166] S5: Install the target dependent component installation package on the basic service image layer to generate a target service image for training the large model, so as to respond to the large model training request.
[0167] An embodiment of the present application also provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in the embodiment of the mirror management method when running, including:
[0168] S1: Construct a basic service image layer, which is used to provide basic components and general environment configurations for supporting large model training;
[0169] S2: Obtain a target large model and at least one dependent component corresponding to the target large model, and generate a one-to-one mapping relationship based on the target large model and the dependent components corresponding to the target large model;
[0170] S3: Based on the one-to-one mapping relationship, determine the installation package of the dependent component corresponding to the target large model, and generate a dependent component resource library based on multiple dependent component installation packages;
[0171] S4: In response to receiving a large model training request, parse the large model training request, and based on the parsing result, extract the target dependent component installation package from the dependent component resource library;
[0172] S5: Install the target dependent component installation package on the basic service image layer to generate a target service image for training the large model, so as to respond to the large model training request.
[0173] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to: various media that can store computer programs such as USB flash drives, read-only memory (ROM for short), random access memory (RAM for short), external hard drives, magnetic disks, or optical discs.
[0174] An embodiment of the present application also provides a computer program product. The computer program product includes a computer program. When the computer program is executed by a processor, it implements the steps in the embodiment of the mirror management method, including:
[0175] S1: Construct a basic service mirror layer, which is used to provide basic components and general environment configurations for supporting large model training;
[0176] S2: Obtain a target large model and at least one dependent component corresponding to the target large model, and generate a one-to-one mapping relationship based on the target large model and the dependent component corresponding to the target large model;
[0177] S3: Based on the one-to-one mapping relationship, determine the dependent component installation packages corresponding to the target large model, and generate a dependent component resource library based on multiple dependent component installation packages;
[0178] S4: In response to receiving a large model training request, parse the large model training request, and based on the parsing result, extract the target dependent component installation package from the dependent component resource library;
[0179] S5: Install the target dependent component installation package on the basic service mirror layer to generate a target service mirror for training the large model, so as to respond to the large model training request.
[0180] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program. When the computer program is executed by a processor, it implements the steps in the embodiment of the mirror management method, including:
[0181] S1: Construct a basic service mirror layer, which is used to provide basic components and general environment configurations for supporting large model training;
[0182] S2: Obtain a target large model and at least one dependent component corresponding to the target large model, and generate a one-to-one mapping relationship based on the target large model and the dependent component corresponding to the target large model;
[0183] S3: Based on the one-to-one mapping relationship, determine the dependent component installation packages corresponding to the target large model, and generate a dependent component resource library based on multiple dependent component installation packages;
[0184] S4: In response to receiving a large model training request, parse the large model training request, and based on the parsing result, extract the target dependent component installation package from the dependent component resource library.
[0185] S5: Install the target dependent component installation package on the basic service image layer to generate a target service image for training the large model, so as to respond to the large model training request.
[0186] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.
[0187] The above has introduced in detail a mirror management method, device, electronic device, and storage medium provided by this application. Specific examples are used in this article to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. A mirror management method, characterized in that: include: Build a basic business image layer, which is used to provide basic components and general environment configuration to support large model training; Acquire a target large model and at least one dependent component corresponding to the target large model, and generate a one-to-one mapping relationship based on the target large model and the dependent component corresponding to the target large model; Based on the one-to-one mapping relationship, determine the dependent component installation package corresponding to the target large model, and generate a dependent component resource library based on a plurality of the dependent component installation packages; In response to receiving the large model training request, parsing the large model training request, and extracting a target dependent component installation package from the dependent component resource library based on the parsing result; The target dependent component installation package is installed on the basic business image layer to generate a target business image for training a large model in response to the large model training request.
2. The image management method according to claim 1, characterized in that: The method further comprises: Obtain the core code logic of the training framework corresponding to the target large model; Based on the core code logic, a training framework project file is generated to construct a training framework project layer.
3. The image management method according to claim 1, characterized in that: Building the basic business image layer includes: Select the target operating system, load the kernel and integrate the hardware drivers; Based on the static compilation mechanism, cross-model common components are preset to generate the basic business image layer.
4. The image management method according to claim 1, characterized in that: Acquiring a target large model and at least one dependent component corresponding to the target large model, and generating a one-to-one mapping relationship based on the target large model and the dependent component corresponding to the target large model includes: Acquire multiple large models supported by the basic business image layer, and a unique identifier corresponding to the target large model; acquire at least one dependent component corresponding to the target large model, and attribute information of the dependent component; Based on the attribute information of the dependent component, generating a dependency item corresponding to the dependent component; Performing one-to-one mapping between the unique identifier corresponding to the target large model and the dependency item corresponding to the dependency component, so as to generate a one-to-one mapping relationship between the target large model and the dependency component corresponding to the target large model; After generating the one-to-one mapping relationship, the method further includes: Obtain the programming language version, deep learning framework version, and distributed training tool version corresponding to the target large model; generate a large model configuration table based on multiple one-to-one mapping relationships, as well as the programming language version, deep learning framework version, and distributed training tool version corresponding to the target large model.
5. The image management method according to claim 1, characterized in that: Generating a dependent component resource library based on the plurality of dependent component installation packages includes: Standardizing the dependent component installation package, wherein the standardized dependent component installation package is accompanied by a metadata file corresponding to the dependent component; The dependent component resource library is generated based on the plurality of dependent component installation packages after standardization.
6. The image management method according to claim 4, characterized in that: Based on the parsing result, extracting the target dependent component installation package from the dependent component resource library includes: Obtaining a parsing result of the large model training request, wherein the parsing result includes at least a large model unique identifier; Based on the large model configuration table, determining the dependency corresponding to the large model unique identifier; Based on the dependency component resource library, the target dependency component installation package corresponding to the dependency item is matched and extracted.
7. The image management method according to claim 6, characterized in that: After extracting the target dependent component installation package from the dependent component resource library, the method further includes: Obtaining a metadata file corresponding to the target dependent component installation package, wherein the metadata file includes at least a compatibility identifier and a hash value; Based on the compatibility identifier and the hash value, the target dependent component installation package is verified.
8. An image management device, characterized in that: include: A basic business image layer construction module, which is used to construct a basic business image layer, and the basic business image layer is used to provide basic components and general environment configurations that support large model training; A large model configuration module, used to obtain a target large model and at least one dependent component corresponding to the target large model, and generate a one-to-one mapping relationship based on the target large model and the dependent component corresponding to the target large model; A dependent component resource library construction module is used to determine the dependent component installation package corresponding to the target large model based on the one-to-one mapping relationship, and generate a dependent component resource library based on a plurality of the dependent component installation packages; A request parsing module, configured to parse the large model training request in response to receiving the large model training request, and extract a target dependent component installation package from the dependent component resource library based on the parsing result; The request response module is used to install the target dependent component installation package on the basic business image layer, generate a target business image for training a large model, and respond to the large model training request.
9. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the image management method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the image management method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Model operation environment determination method and device, storage medium and electronic equipment
CN118227212A
Dynamic dependency driven software test task organization system and organization method
CN118227471A
Container mirror image creating method and device, electronic equipment and storage medium
CN118444943A