A Method for Removing Redundancy of Container Images by Reconstructing Mirror Layers
Through the mirror layer reconstruction method, the redundancy of the container image layer is identified and optimized, and divided into unique and shared layers through parallel traversal and threshold judgment, which solves the redundancy problem of container image layer and improves the image usage efficiency and compatibility.
Patent Information
- Application Number
- CN202310702283.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-14
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-06-14
AI Technical Summary
In the prior art, the layer division of container images is based on Dockerfile instructions, resulting in high redundancy between layers, and it is impossible to effectively share files between different layers, resulting in high storage and transmission costs.
By traversing the mirror layer files in parallel, metadata is generated and merged views are established, redundant files are identified, divided into unique layers and shared layers, and the threshold is used to determine whether to perform mirror reconstruction, and the layer structure is optimized to reduce redundancy.
Without destroying the native layer structure, reduce mirror redundancy, improve mirror usage efficiency, reduce storage and transmission costs, and maintain mirror compatibility.
Smart Images

Figure CN116841972B_ABST
Abstract
Description
Technical Field:
[0001] The present invention belongs to the technical field of containerized services, and particularly relates to a method for removing redundancy of container images through mirror layer reconstruction. Background Art:
[0002] With the development and increasingly widespread application of cloud computing technology, containerization technology has become increasingly important. In containerization technology, container images are a very important concept, which provides a lightweight, portable, and reusable way to build and deploy application programs. However, since the size of container images is usually relatively large, this will lead to higher storage and transmission costs, and also take more time and resources during the deployment process. Therefore, how to optimize the deployment cost of container images has become an important issue. The background art mainly involved in the present invention includes the following aspects:
[0003] Container: A container is a lightweight virtualization technology that allows application programs to run in an isolated environment without being affected by the host operating system. A container is a virtualization technology at the operating system level. Different from traditional virtual machines, it does not require an additional virtual machine manager, nor does it need to virtualize the hardware. Therefore, the performance of containers is very high. Containers play a very important role in application development and deployment. They provide a lightweight, portable, and reusable way to build and deploy application programs. Using containers, developers can package application programs and all dependencies into an image, and this image can run on any platform that supports containerization technology without additional configuration and installation. The isolation of containers enables application programs to run in an independent environment, which means that different application programs can run on the same host without interfering with each other. In addition, containers can also be connected to other containers and external services through the network to achieve complex application architectures. In addition to developing and deploying application programs, containers can also be used in fields such as testing, continuous integration, and delivery. Containers can make the testing and deployment of application programs more automated and repeatable, while also improving the efficiency and productivity of developers.
[0004] Container Image: A container image is a read-only file that contains all the files and dependencies required to build and run a container. Container images typically include an operating system, applications, libraries, configuration files, and other dependencies. They can be provided by developers, system administrators, or third parties and are usually obtained through image repositories or other distribution channels. A container image is the basis for a container. The container image contains a set of layers, each of which is a read-only file system. When the container runs, these layers are combined to form a complete file system, and the application is run on this basis. Since each layer of the image is read-only, containers can isolate different applications and runtime environments to ensure security and reliability.
[0005] Image Layer: A container image consists of multiple image layers, each of which is a read-only file system. Image layers are a fundamental part of a container image. They contain applications, libraries, configuration files, dependencies, etc. The image layers are combined through the union mount technology, which can mount multiple image layers into a file system (container) to form a complete container. The advantage of image layers is that they can improve the repeatability and portability of container images. Each image layer is an independent component that can be shared and reused among different images. This can reduce the size of the image, thereby improving the efficiency of transmission and deployment. In addition, the immutability and isolation of image layers can ensure the security and reliability of container images.
[0006] The resource isolation ability of container technology enables multiple containers to run on the same server without interfering with each other. However, this isolation also hinders file reuse between images, resulting in file redundancy. Currently, the division of layers in the image is based on the Dockerfile (configuration file), that is, each line of instructions in the Dockerfile constructs an image layer, and sharing can only occur when two layers are exactly the same. However, the randomness of Dockerfile instructions hinders data sharing, resulting in a high degree of redundancy between different layers. Specifically, it results in many similar but different layers, that is, some files in the two layers are exactly the same, but there are also individual files that are different. However, the current reuse mechanism of Docker can only share two completely identical layers. Even if only one file is different between the two layers, the entire layer cannot be shared. Therefore, it is necessary to reconstruct the files contained in the image layer through this solution to make as many layers as possible completely identical, thereby reducing image redundancy without destroying the native layer structure and maintaining high compatibility. Summary of the Invention:
[0007] Aiming at the technical problems existing in the prior art, the present invention provides a method for reducing the file redundancy between layers by reconstructing the files included in different layers of the image, while also maintaining compatibility with the native image layer structure; the present invention reconstructs the files included in the image layer to make as many layers as possible completely identical, so as to reduce the image redundancy on the premise of not destroying the native layer structure and maintaining high compatibility.
[0008] The core content of the invention can be summarized as follows:
[0009] A method for removing redundancy of container images by reconstructing image layers, comprising the following steps:
[0010] S1. Adopt a parallel traversal method to collect the path, name, size and hash value of all files in the image layer to obtain the corresponding image file metadata;
[0011] S2. Establish an image merge view according to the image file metadata;
[0012] S3. Compare the merge views of different images to determine the redundancy of files;
[0013] S4. Divide the image unique layer and the image shared layer according to the redundancy of the image files; wherein:
[0014] Files with the same file name and hash value that appear in different images are divided into the shared layer;
[0015] Files that exist only in one image are divided into the unique layer;
[0016] S5. Determine whether to perform image reconstruction through a trade-off threshold; wherein:
[0017] Calculate whether the size of the shared layer files exceeds the layer threshold; if satisfied, create the shared layer;
[0018] Otherwise, cancel the creation of the shared layer, and divide the files of this layer into the unique layer;
[0019] Calculate whether the total size of all shared layers newly added by image reconstruction exceeds the trade-off threshold. If satisfied, perform image reconstruction; otherwise, return to step S4.
[0020] Further, the process of establishing an image merge view according to the image file metadata:
[0021] Merge the files with different paths and file names in the lower-layer image into the upper-layer image;
[0022] Overwrite the files with the same path and file name in the upper-layer image on the lower-layer image.
[0023] Delete the hidden files in the lower-layer image.
[0024] Beneficial effects
[0025] Advantages of the present invention compared with the prior art:
[0026] In the present invention, mirror reconstruction can improve the mirror usage efficiency by adjusting the number of mirror layers and the files included in each layer, so that more completely identical layers appear, thereby facilitating layer sharing between different mirrors and reducing redundant files between mirrors. Therefore, mirror reconstruction not only reduces the mirror storage space in the mirror repository, but also benefits the container deployment process by avoiding redundant file transfer and extraction, while retaining the native layer structure design and compatibility. Description of the drawings:
[0027] Figure 1 is the flow chart of container mirror reconstruction in the present invention;
[0028] Figure 2 is an example diagram of container mirror reconstruction related to the present invention. Detailed implementation manners:
[0029] The following will combine the attached Figures 1 - 2 to make the following description of the present invention:
[0030] The goal of mirror reconstruction is to enhance the mirror layer reuse function and achieve near file-level redundancy removal ability. The overall process of mirror reconstruction is as Figure 1 shown, including three main steps:
[0031] Step 1: Generate file metadata for each mirror according to its initial structure, and then create a mirror merge view by merging mirror layers.
[0032] Step 2: Use the merge view to identify redundant files between mirrors.
[0033] Step 3: Divide files into unique layers and shared layers according to the redundancy information.
[0034] To optimize and accelerate this process, steps 1 and 2 are only inferred based on file metadata, and the actual file operations will only be executed after the threshold conditions in step 3 are met. The following will be described in detail with the Figure 2 example.
[0035] Step 1: Generate file metadata and merge view. In this step, all files in the mirror will be traversed in parallel and the path, name, size, and hash value (SHA 256) will be collected to generate file metadata for each mirror. Then, based on the metadata, layer-by-layer merging will be performed according to the following method to create a merge view of the mirror:
[0036] (i) If the files and folders in the lower layer have different paths and file names from those in the upper layer, they will be merged into the upper layer, as shown by file x in Figure 2 when mirror 1 is merged based on layer 1 and layer 2.
[0037] (ii) If the files and folders in the lower layer have the same paths and file names as those in the upper layer, regardless of whether their contents are repeated, the upper layer files and folders will overwrite the lower layer, as shown by file f in Figure 2 .
[0038] (iii) Delete the files and folders marked by the "whiteouts" mechanism, such as.wh..wh..opq (hiding all sub-files) and.wh.z (hiding file z) in Figure 2 .
[0039] Step 2: Determine the shareability of each file according to the merged view. First, generate a global key-value table by traversing the merged views of all mirrors in parallel:
[0040] The key is the hash value of the file, and the value is the list of mirrors containing the file. Files with more than one mirror in the mirror list are potential shareable files and are classified into the shared layer, such as file f and file i in Figure 2 . In addition, files that appear in only one mirror are retained in the unique layer, such as file x and file q in Figure 2 . Finally, the new layer structure of each mirror can be inferred from the file metadata.
[0041] Step 3: Reconstruct the new layer structure. To optimize efficiency, two thresholds are set (users can customize the threshold sizes according to their needs):
[0042] (i) Layer creation threshold:
[0043] To avoid generating too many layers, a shared layer will only be created when the size of the shared layer exceeds the threshold; otherwise, the creation of the shared layer will be cancelled and the files will be retained in the unique layer.
[0044] (ii) Trade-off threshold:
[0045] Considering the cost of mirror reconstruction, it is necessary to calculate the size of the reusable files increased by this reconstruction. If the size exceeds the threshold, the reconstruction process will be executed; otherwise, this reconstruction will be abandoned. When the reconstruction is executed to meet the threshold requirements, the files will be re-partitioned according to the layer structure planned in step 2. As shown in Figure 2 , file f and file i will be moved to the specified shared path / a / z / , and their original positions will be replaced with soft links pointing to the files in / a / z / .
Claims
1. A method for removing redundancy from container images by reconstructing mirror layers, characterized in that, It includes the following steps: S1. Use a parallel traversal method to collect the path, name, size, and hash value of all files in the image layer to obtain the corresponding image file metadata; S2. Establish an image merge view based on the image file metadata; S3. Compare the merge views of different images to determine the redundancy of files; S4. Divide the image into a unique layer and a shared layer according to the redundancy of the image files; where: Files with the same file name and hash value that appear in different images are divided into the shared layer; Files that exist only in one image are divided into the unique layer; S5. Determine whether to perform image reconstruction through a trade-off threshold; where: Calculate whether the size of the files in the shared layer exceeds the layer threshold; if satisfied, create the shared layer; otherwise, cancel the creation of the shared layer, and the files in this layer are divided into the unique layer; Calculate whether the total size of all newly added shared layers during image reconstruction exceeds the trade-off threshold. If satisfied, perform image reconstruction; otherwise, return to step S4.
2. The method for removing redundancy of container images by mirror layer reconstruction according to claim 1, characterized in that The process of establishing an image merge view according to the image file metadata: Merge the files with different paths and file names in the lower-layer image into the upper-layer image; Overwrite the files with the same path and file name in the upper-layer image on the lower-layer image; Delete the hidden files in the lower-layer image.
Citation Information
Patent Citations
Mirror image processing method and device, electronic equipment and computer readable storage medium
CN113934510A