Method and system for achieving mirror dynamic caching based on incremental snapshots
By using incremental snapshot technology and ARIMA model to predict cache requirements in an edge computing environment, dynamic image caching is realized, solving the problem of bottlenecks in the central mirror warehouse and the repetitive occupation of traditional cache methods, and improving system efficiency and scalability.
Patent Information
- Application Number
- PCT/CN2024/135805
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-12
- Filing Date
- 2024-11-29
- Publication Date
- 2025-06-19
AI Technical Summary
In edge computing scenarios, the central mirror warehouse becomes a bottleneck, and the traditional mirror caching method causes repeated space and bandwidth consumption, and requires a lot of manpower to pre-cache.
The image dynamic cache method based on incremental snapshots is adopted to save the image layered through the COW file system, and the image is cached only when the basic image is not cached. Otherwise, only the incremental part is cached, and the cache needs are predicted through the ARIMA model to automatically trigger the cache.
It reduces the use of network bandwidth and mirror cache space, reduces the demand for manual operations, improves the startup speed of virtual machines, optimizes resource utilization of edge clusters, and improves the scalability of the system.
Smart Images

Figure CN2024135805_19062025_PF_FP_ABST
Abstract
Description
A method and system for implementing dynamic mirror caching based on incremental snapshots Technical Field
[0001] The present invention relates to the technical field of dynamic mirror caching, and in particular to a method and system for implementing dynamic mirror caching based on incremental snapshots. Background Art
[0002] In edge computing scenarios, creating a virtual machine in an edge cluster often requires pulling the images required by the virtual machine from a central image repository, or caching commonly used images from the central image repository to the edge cluster in advance. However, both methods require obtaining images from the central image repository. As the number of edge clusters increases, the central image repository will become a bottleneck.
[0003] As the business is used, more and more images may need to be cached, and manually pre-caching the images in advance would also consume too much manpower.
[0004] The increase in cached images consumes increasing space in edge clusters. However, these new images are often based on the same kernel image, with various software and environments deployed on top of it before being packaged into new images. Traditional image caching methods simply cache the entire image at the edge, causing the base image to be cached repeatedly. This not only takes up space but also consumes significant bandwidth and requires a long cache time. Summary of the Invention
[0005] The purpose of the present invention is to provide a method and system for implementing dynamic mirror caching based on incremental snapshots, so as to reduce the occupation of network bandwidth and mirror cache space.
[0006] In one aspect, the present invention proposes a method for implementing dynamic image caching based on incremental snapshots, comprising:
[0007] S1: Store the base image in layers in the image repository on the central cloud.
[0008] S2: caching the base image in a designated edge cluster based on the first cache triggering mode or the second cache triggering mode;
[0009] The first cache triggering mode is to specify a to-be-cached image, and cache the to-be-cached image directly from the central cloud to the specified edge cluster;
[0010] The second cache triggering mode is to periodically judge whether cache is needed. If the judgment result is yes, cache is started; if the judgment result is no, the process waits to enter the next cycle.
[0011] Furthermore, the second cache triggering method includes: presetting a cycle period, in each of the cycle periods, if caching is not required, waiting to enter the next cycle period; if caching is required, determining whether it has been cached, if the determination result is yes, waiting to enter the next cycle period, if the determination result is no, triggering the mirror cache logic and waiting to enter the next cycle period.
[0012] Furthermore, the image caching logic includes: specifying the image to be cached, and determining whether the image to be cached has been cached in the specified edge cluster; if the determination result is yes, exiting the second cache triggering mode; if the determination result is no, querying whether the image to be cached has been cached in other edge clusters; if the determination result is yes, triggering edge-to-edge caching; if the determination result is no, triggering cloud-edge caching.
[0013] Furthermore, the edge-to-edge cache is a cache form between the edge clusters, and the cloud-edge cache is a cache form from the central cloud to the edge cluster; both the cloud-edge cache and the edge-to-edge cache are cached based on incremental snapshots, and verification is performed before transmitting the image. If the base image has been cached, only the incremental part is cached.
[0014] Furthermore, the cloud-edge caching includes: the central cloud determines based on the image cache information that the image to be cached has not been cached in other clusters, the central cloud exports the image in the image warehouse, and communicates to the designated edge cluster, and the designated edge cluster caches the image to local storage after receiving it.
[0015] Furthermore, the edge-to-edge caching includes: the central cloud determines based on the image cache information that the image to be cached has been cached in other clusters, then obtains cluster load information and network delay, and based on the cluster load information and the network delay, obtains the optimal cluster through a scheduling algorithm, and the cluster to be cached obtains the image from the optimal cluster and caches it.
[0016] Furthermore, the scheduling algorithm includes: detecting load information of each edge cluster, the load information including CPU usage, memory usage, and disk space usage, and evaluating the load of each edge cluster based on the CPU usage, the memory usage, and the disk space usage;
[0017] detecting a network delay of each edge cluster, where the network delay is used to determine a distance between any two edge clusters;
[0018] Record the image list that has been cached by each edge cluster, and obtain the image status that has been cached by each edge cluster.
[0019] Furthermore, the incremental snapshot includes: creating a snapshot layer for the image through the COW file system, the edge cluster exports the incremental snapshot image for transmission, and the designated cluster receives and stores the incremental snapshot image.
[0020] Furthermore, determining whether caching is required in each of the cycles includes:
[0021] Collect historical data on the cache status and image usage of each edge cluster and convert it into a time series.
[0022] The preset Ci,j,t represents the cache status of image i on edge cluster j at time t;
[0023] The preset Ui,j,t represents the usage of image i on edge cluster j at time t; the image cache status on each edge cluster is taken as a time series Sj,t
[0024] Where nj is the number of mirrors in edge cluster j, which is n;
[0025] Each of the time series is modeled using the ARIMA model to predict future time series values.
[0026] Compared with the prior art, the present invention has the following beneficial effects:
[0027] Through model prediction and automated caching, the system will reduce the need for manual operations, reduce manual intervention and maintenance of image caches, thereby saving time, cost and human resources; automatic caching strategies can ensure that the required images are in place in the edge cluster in advance, reducing the waiting time for pulling images from the central image warehouse when the virtual machine starts, and accelerating the startup speed of the virtual machine; by sharing images between edge clusters, the limited bandwidth resources in the edge computing environment are rationally utilized, the network traffic to the central image warehouse is reduced, and the network bandwidth usage is reduced; through incremental snapshots and COW file systems, only incremental data needs to be stored instead of the entire image, which reduces the storage space usage and ensures the effective use of storage resources; by supporting pulling images from nearby edge clusters and choosing to reasonably distribute images between different clusters, the load is dispersed and the pressure on the central image warehouse is reduced, which helps to ensure the image acquisition process of each edge cluster and improves the scalability of the system.
[0028] On the other hand, the present invention also provides a system for implementing dynamic image caching based on incremental snapshots, which is applied to the above-mentioned method for implementing dynamic image caching based on incremental snapshots. The system includes: a central cloud and several edge clusters, wherein the central cloud is deployed with an image-cloud module and a central storage module, and each edge cluster is deployed with an image-edge module and an edge storage module;
[0029] The image-cloud module is used to:
[0030] (1) Establishing grpc streaming bidirectional communication with the image-edge module;
[0031] (2) receiving information reported by the image-edge module for policy determination;
[0032] (3) Recording image cache information for model generation;
[0033] (4) Determine whether the edge cluster needs to cache an image based on the model;
[0034] (5) Determining the image cache mode based on the image cache information;
[0035] (6) Support sending images to the image-edge module;
[0036] The image-edge module is used to:
[0037] (1) regularly updating the information of the edge cluster from the image-cloud module;
[0038] (2) reporting the load information of the edge cluster to the image-cloud module of the other edge clusters;
[0039] (3) used to receive or send images from other image-edge modules; used to receive images from the image-cloud module;
[0040] (4) Use network measurement tools to detect the network delay from the current cluster to other clusters;
[0041] (5) Collect the load information of the current edge cluster.
[0042] Furthermore, it also includes:
[0043] If a new edge cluster is added, the central cloud updates the new edge cluster topology information to each edge cluster;
[0044] The image-edge module in the central cloud regularly measures and records the network status of the edge cluster;
[0045] The image-edge module regularly collects information about the edge cluster, including CPU usage, memory usage, and disk space usage. After collecting the information, it reports the network status and cache status to the image-cloud module.
[0046] The image-cloud module collects information of each edge cluster and stores it for analysis.
[0047] It is understandable that the above-mentioned control method and system for recycling production waste have the same beneficial effects and will not be described in detail here. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0049] 1 is a framework diagram of a system for implementing dynamic mirror caching based on incremental snapshots according to an embodiment of the present invention;
[0050] 2 is a schematic diagram of a hierarchical storage form of a mirror dynamic cache system based on incremental snapshots according to an embodiment of the present invention;
[0051] 3 is a flow chart of a second cache triggering method according to an embodiment of the present invention;
[0052] 4 is a flow chart of determining mirror cache logic according to an embodiment of the present invention;
[0053] FIG5 is a diagram showing the operation steps of edge-to-edge caching according to an embodiment of the present invention;
[0054] FIG6 is a diagram showing the operation steps of cloud-edge caching according to an embodiment of the present invention;
[0055] FIG7 is a diagram of the operation steps when a new edge cluster is added according to an embodiment of the present invention.
[0056] In the figure, 1, central cloud; 2, edge cluster; 21, first edge cluster; 22, second edge cluster; 23, third edge cluster. DETAILED DESCRIPTION
[0057] The following is a clear and complete description of the technical solutions in the embodiments of the present invention. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of the present invention.
[0058] In the description of this application, it should be understood that the terms "center", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc., indicating the orientation or position relationship, are based on the orientation or position relationship shown in the accompanying drawings, and are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on this application.
[0059] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature specified as "first" or "second" may explicitly or implicitly include one or more of such features. Throughout this application, unless otherwise specified, "plurality" means two or more.
[0060] In the description of this application, it should be noted that, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense. For example, they can refer to fixed connections, detachable connections, or integral connections; mechanical connections or electrical connections; direct connections or indirect connections through an intermediate medium; and internal connections between two components. Those skilled in the art will understand the specific meanings of the above terms in this application based on the specific circumstances.
[0061] 1-7 , a method for implementing dynamic image caching based on incremental snapshots according to a preferred embodiment of the present invention includes:
[0062] S1: The base image is stored in layers in the image repository located in the central cloud 1;
[0063] S2: Based on the first cache triggering mode or the second cache triggering mode, cache the base image to the specified edge cluster 2;
[0064] Among them, the first cache triggering mode is to specify the image to be cached, and cache the image to be cached directly from the central cloud 1 to the specified edge cluster 2;
[0065] The second cache triggering mode is to periodically judge whether caching is required. If the judgment result is yes, caching is started. If the judgment result is no, the process waits to enter the next cycle.
[0066] In this embodiment, a COW (Copy-on-write) file system is used for layered storage. When caching a new image, if the base image has already been cached, only the incremental portion needs to be cached, as shown in Figure 2. Specifically, in this embodiment, Ceph is used as the storage module. The image repository of central cloud 1 uses Ceph as the storage module, and the storage mode of edge cluster 2 also uses Ceph. The data of central cloud 1 is stored in MySQL, and the image-cloud module and image-edge module are implemented in Golang.
[0067] The image-cloud module is deployed in the central cloud 1, and the image-edge module is deployed in each edge cluster 2. The image-cloud module collects and stores the edge cluster 2 load information reported by the image-edge module in each edge cluster 2 and the network status information between edge clusters 2 in the MySQL database for use in the DWLB scheduling algorithm. The image-cloud module stores the image cache information in MySQL and uses it as raw data to generate the ARIMA model. Using Ceph storage as the image repository, the following commands are used to implement incremental export: (1) Create a snapshot layer for the image using the command rbd snap create; (2) Export the incremental snapshot using the command rbd export-diff; (3) Import the image cache using the command rbd import-diff.
[0068] The first cache triggering method can be used for manual triggering, and the operator caches the specified image in the specified edge cluster 2; the second cache triggering method can be used for automatic triggering. The image-cloud module determines whether a certain edge cluster 2 needs to cache a certain image at a certain point in time based on model prediction. If so, the image caching logic is triggered.
[0069] It should be emphasized that ceph is used in the above embodiment. In addition, it also includes but is not limited to zfs or btrfs, which support features such as image layering, image incremental snapshots, and incremental image export and import, and can be used as a warehouse for image cache.
[0070] As shown in FIG3 , in some embodiments, the second cache triggering method includes: presetting a cycle period, and in each cycle period, if caching is not required, waiting to enter the next cycle period; if caching is required, determining whether it has been cached, and if the determination result is yes, waiting to enter the next cycle period; if the determination result is no, triggering the mirror cache logic and waiting to enter the next cycle period.
[0071] As shown in Figure 4, in some embodiments, the image caching logic includes: specifying the image to be cached, and determining whether the image to be cached has been cached in the specified edge cluster 2; if the judgment result is yes, exiting the second cache triggering mode; if the judgment result is no, querying whether the image to be cached has been cached in other edge clusters 2; if the judgment result is yes, triggering edge-to-edge caching; if the judgment result is no, triggering cloud-edge caching.
[0072] It can be understood that through the preset cycle, the system can periodically check whether caching operations are needed, which helps to optimize the caching strategy and ensure that the cached image data is always up to date; when it is determined that caching is needed, the cache trigger is automatically executed, reducing the need for manual intervention, thereby reducing management and maintenance costs; at the same time, periodic judgment can reduce the continuous occupation of network bandwidth and trigger caching operations only when necessary.
[0073] It should be noted that the preset cycle period can be selected according to actual conditions, for example, the cycle period can be updated and changed in real time according to data on network bandwidth usage and image cache space usage.
[0074] In addition, the method and system for implementing dynamic image caching based on incremental snapshots are valuable in a variety of scenarios: for edge computing, pushing computing and data storage close to the data generation source, this method can help the edge cluster 2 cache images locally, thereby reducing dependence on the central cloud image warehouse and improving service response speed; for cloud service providers, being able to more effectively manage and optimize image caches can reduce storage and bandwidth costs and provide better service quality; in a virtualized environment, quickly starting virtual machines is critical to performance and user experience. By automating image caching, the virtual machine startup time can be significantly reduced and the efficiency of the virtualized environment can be improved; for containerized applications, which have become the mainstream way of deploying modern applications, in a container orchestration system, the method of this embodiment can be applied to start container instances faster and reduce network traffic; in cases where similar virtual machines or container instances need to be deployed on a large scale, automated incremental snapshot caching can significantly improve efficiency and reduce network transmission and storage resource usage; in cases where data distribution is required between multiple locations or edge locations, incremental snapshot caching can reduce the amount of data transmitted across the network, thereby saving bandwidth costs.
[0075] In some embodiments, edge-to-edge caching is a caching form between edge clusters 2, and cloud-edge caching is a caching form from the central cloud 1 to the edge cluster 2; both cloud-edge caching and edge-to-edge caching are based on incremental snapshots, and verification is performed before transmitting the image. If the base image has been cached, only the incremental part is cached.
[0076] Specifically, as shown in Figure 5, the image-cloud module, based on image cache information, presumably obtains that the image is already cached in the first edge cluster 21 and the second edge cluster 22. It then uses the scheduling algorithm to make a comprehensive judgment based on information such as the load information and network status of edge cluster 2. If the first edge cluster 21 is determined to be the optimal cluster, it sends a command to the image-edge module of the third edge cluster 23, instructing it to obtain the image from the first edge cluster 21 for caching.
[0077] In some embodiments, cloud-edge caching includes: the central cloud 1 determines based on the image cache information that the image to be cached has not been cached in other clusters, the central cloud 1 exports the image in the image warehouse, communicates to the designated edge cluster 2, and the designated edge cluster 2 caches the image to local storage after receiving it.
[0078] It should be noted that, as shown in Figure 6, the image-cloud module of the central cloud 1 determines whether the image has been cached in other clusters based on the image cache information. If not, the image-cloud module exports the image from the central image repository and communicates with the image-edge of the specified edge cluster 2 through grpc bidirectional streaming. After receiving the image, the image-edge caches the image to local storage.
[0079] In some embodiments, edge-to-edge caching includes: the central cloud 1 determines, based on the image cache information, that the image to be cached has been cached in other clusters, then obtains cluster load information and network delay, and based on the cluster load information and network delay, obtains the optimal cluster through a scheduling algorithm, and the cluster to be cached obtains and caches the image from the optimal cluster.
[0080] In some embodiments, the scheduling algorithm includes: detecting load information of each edge cluster 2, the load information including: CPU usage, memory usage, and disk space usage, and evaluating the load of each edge cluster 2 according to the CPU usage, memory usage, and disk space usage;
[0081] Detect the network delay of each edge cluster 2. The network delay is used to determine the distance between any two edge clusters 2.
[0082] Record the image list that has been cached by each edge cluster 2, and obtain the image status of each edge cluster 2 that has been cached.
[0083] It should be noted that network latency between edge clusters 2 can be measured using network measurement technologies such as Ping and traceroute. This latency information can then be used to determine the distance between each edge cluster 2 and the other clusters. Monitoring tools can be used to monitor the load of each edge cluster 2, including CPU usage, memory usage, and disk space usage.
[0084] It should be noted that the scheduling algorithm can be:
[0085] Use the top command to view CPU usage, free -h to view memory usage, and df -h to view disk space usage. These load indicators are reported to the cloud and saved in the database.
[0086] The distance-weighted load balancing index DWLB is preset and satisfies the following relationship: DWLB(i)=α×Load(i)+(1-α)×Distance(i, j);
[0087] Where DWLB(i) represents the index value of caching the image to the i-th edge cluster 2; Load(i) represents the load of the i-th cluster; Distance(i, j) represents the distance between the i-th cluster and the j-th cluster; α represents the weight coefficient;
[0088] Regardless of cloud-edge caching or edge-edge caching, after caching is successful, the edge image-edge module will report the cached image list of this set of edge clusters to the cloud, and the cloud will record the cached image list of each edge cluster through the database.
[0089] Using the cache list, we determine which edge clusters have already cached the image. These clusters are considered candidate clusters. Next, based on the load (Load(i)) of the target cluster (the cluster where the image is to be cached), we consider factors such as low CPU and memory usage, as well as sufficient storage space. If there is insufficient remaining space, the image cannot be cached. If there is sufficient space, we calculate the distance from the target cluster to the candidate clusters (low network latency and minimal packet loss are considered optimal). Using the formula DWLB(i), we determine a relatively optimal cluster from the candidate clusters for caching.
[0090] In some embodiments, the incremental snapshot includes: creating a snapshot layer for the image through the COW file system, the edge cluster 2 exporting the incremental snapshot image for transmission, and the designated cluster receiving and storing the incremental snapshot image.
[0091] In some embodiments, determining whether caching is required in each cycle includes:
[0092] Collect historical data on cache status and image usage of each edge cluster 2 and convert it into time series.
[0093] The preset Ci,j,t represents the cache status of image i on edge cluster j at time t;
[0094] The preset Ui,j,t represents the usage of image i on edge cluster j at time t;
[0095] Where nj is the number of mirrors in edge cluster j, which is n;
[0096] Model each time series using the ARIMA model to predict future time series values.
[0097] Specifically, the prediction model is generated to achieve prediction based on statistical historical cache records, using the ARIMA model in the time series model, which contains autoregressive terms, difference terms and moving average terms to predict future time series values.
[0098] The image cache status on each edge cluster 2 is treated as a time series and modeled. Historical data on the cache status (whether there is a cache, cache size, etc.) and image usage (image name, number of uses, usage duration, etc.) of each edge cluster 2 is collected and converted into a time series.
[0099] Assume that Ci,j,t represents the cache status of image i on edge cluster j at time t, and Ui,j,t represents the usage of image i on edge cluster j at time t. Then, the image cache status on each edge cluster 2 can be regarded as a time series Sj,t, and satisfies the following relationship:
[0100] Finally, an ARIMA model is used to model each time series and predict future time series values.
[0101] It should be noted that by collecting and analyzing historical cache status and image usage data from edge cluster 1, and using the ARIMA model for modeling and prediction, it is possible to more accurately determine whether image caching is needed in each cycle. This can significantly reduce unnecessary caching operations, thereby reducing the waste of storage and computing resources. By predicting future image caching needs, the system can avoid unnecessary data transmission and storage, thereby reducing network bandwidth usage and storage space waste. This helps improve overall system efficiency and reduce operating costs. The ARIMA model allows the system to make real-time decisions in each cycle, meaning it can quickly adapt to changing needs and conditions, ensuring that the image cache in the edge cluster remains up-to-date and efficient. Using the ARIMA model for prediction and decision-making reduces the need for manual intervention. The system can automatically make caching decisions based on historical data and model output, reducing human error and costs. The ARIMA model generates predictions based on historical data, making cache management more data-driven, with decisions based on actual usage and trends rather than static rules or routine operations.
[0102] In addition, an embodiment of the present invention further provides a system for implementing dynamic image caching based on incremental snapshots, which is applied to the above-mentioned method for implementing dynamic image caching based on incremental snapshots. The system includes: a central cloud 1 and a plurality of edge clusters 2. In this embodiment, the number of edge clusters is three;
[0103] The central cloud 1 is deployed with an image-cloud module and a central storage module, and each edge cluster 2 is deployed with an image-edge module and an edge storage module;
[0104] The image-cloud module is used to:
[0105] (1) Establish grpc streaming bidirectional communication with the image-edge module;
[0106] (2) Receive information reported by the image-edge module for policy judgment;
[0107] (3) Recording image cache information for model generation;
[0108] (4) Based on the model, determine whether edge cluster 2 needs to cache the image;
[0109] (5) Determining the image cache mode based on the image cache information;
[0110] (6) Support sending images to image-edge module;
[0111] The image-edge module is used to:
[0112] (1) Update the information of edge cluster 2 from the image-cloud module regularly;
[0113] (2) Report the load information of edge cluster 2 to the image-cloud module of other edge clusters 2;
[0114] (3) Used to receive or send images from other image-edge modules; used to receive images from the image-cloud module;
[0115] (4) Use network measurement tools to detect the network delay from the current cluster to other clusters;
[0116] (5) Collect the load information of the current edge cluster 2.
[0117] In some embodiments, as shown in FIG7 , the following further comprises:
[0118] If a new edge cluster is added, the central cloud 1 updates the new edge cluster topology information to each edge cluster 2;
[0119] The image-edge module of the central cloud 1 regularly measures and records the network status of the edge cluster 2;
[0120] The image-edge module periodically collects information about edge cluster 2, including CPU usage, memory usage, and disk space usage. After collecting this information, it reports the network and cache status to the image-cloud module.
[0121] The image-cloud module collects information from each edge cluster 2 and stores it in the central storage module for analysis.
[0122] It should be noted that updating the topology information of the new edge cluster to each edge cluster helps maintain the accuracy of the network topology and ensures the correct transmission and caching of data. This is especially important for environments with multiple edge clusters 2, because the addition of new edge clusters may cause changes in the network topology. The image-edge module of the central cloud 1 regularly measures and records the network status of the edge cluster 2, which is beneficial for understanding the network status and can help select the best edge cluster for caching to improve performance and service quality, monitor network conditions and take appropriate measures. The image-edge module regularly collects information about the edge cluster 2, including CPU usage, memory usage and disk space usage, which is beneficial. Understanding the resource utilization of each edge cluster 2 can better allocate caching tasks and achieve load balancing; based on the resource utilization information, the central cloud 1 can take corresponding measures, such as dynamically adjusting the allocated resources or providing additional resources, to optimize the performance of the edge cluster; the image-edge module reports the network status and cache status to the image-cloud module, which is beneficial for the image-cloud module to use the reported information to make more intelligent decisions, select the best edge cluster for caching or adjust the caching strategy; the image-cloud module can monitor the network status and resource utilization of each edge cluster 2 in real time to ensure that the performance reaches the required level.
[0123] The above description is only an embodiment of the present invention, but it cannot be used to limit the scope of the present invention. Any structural changes made according to the present invention should be deemed to fall within the scope of protection of the present invention and be subject to restrictions as long as they do not lose the essence of the present invention.
[0124] It should be noted that the system provided in the above embodiment is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the modules or steps in the embodiments of the present invention can be further decomposed or combined. For example, the modules in the above embodiment can be combined into one module, or further divided into multiple sub-modules to complete all or part of the functions described above. The names of the modules and steps involved in the embodiments of the present invention are only for distinguishing the modules or steps and are not to be regarded as improper limitations of the present invention.
[0125] The term "comprise" or any other similar term is intended to cover non-exclusive inclusion such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed or inherent to such process, method, article, or apparatus.
[0126] Thus far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art may make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will fall within the scope of protection of the present invention.
[0127] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention.
Claims
1. A method for implementing dynamic image caching based on incremental snapshots, characterized in that: include: S1: The basic images are stored in layers in the image repository located in the central cloud; S2: Based on the first cache triggering mode or the second cache triggering mode, cache the base image in a designated edge cluster; The first cache triggering mode is to specify a to-be-cached image, and cache the to-be-cached image directly from the central cloud to the specified edge cluster; The second cache triggering mode is to periodically determine whether cache is required. If the determination result is yes, cache is started; if the determination result is no, the next cycle is waited for.
2. The method for implementing dynamic image caching based on incremental snapshots according to claim 1, characterized in that: The second cache triggering method includes: presetting a cycle period, in each cycle period, if caching is not required, waiting to enter the next cycle period; if caching is required, determining whether it has been cached, if the determination result is yes, waiting to enter the next cycle period, if the determination result is no, triggering the mirror cache logic and waiting to enter the next cycle period.
3. The method for implementing dynamic image caching based on incremental snapshots according to claim 2, characterized in that: The image caching logic includes: specifying the image to be cached, and determining whether the image to be cached has been cached in the specified edge cluster; if the determination result is yes, exiting the second cache triggering mode; if the determination result is no, querying whether the image to be cached has been cached in other edge clusters; if the determination result is yes, triggering edge-to-edge caching; if the determination result is no, triggering cloud-to-edge caching.
4. The method for implementing dynamic image caching based on incremental snapshots according to claim 3, characterized in that: The edge-to-edge cache is a cache form between the edge clusters, and the cloud-edge cache is a cache form from the central cloud to the edge cluster; both the cloud-edge cache and the edge-to-edge cache are cached based on incremental snapshots, and a check is performed before the image is transmitted. If the base image has been cached, only the incremental part is cached.
5. The method for implementing dynamic image caching based on incremental snapshots according to claim 4, characterized in that: The cloud-edge caching includes: the central cloud determines that the image to be cached has not been cached in other clusters based on the image cache information, the central cloud exports the image in the image warehouse, communicates to the designated edge cluster, and the designated edge cluster caches the image to local storage after receiving it.
6. The method for implementing dynamic image caching based on incremental snapshots according to claim 4, characterized in that: The edge-to-edge caching includes: the central cloud determines, based on the image cache information, that the image to be cached has been cached in other clusters, then obtains cluster load information and network delay, and based on the cluster load information and the network delay, obtains the optimal cluster through a scheduling algorithm, and the cluster to be cached obtains and caches the image from the optimal cluster.
7. The method for implementing dynamic image caching based on incremental snapshots according to claim 6, characterized in that: The scheduling algorithm includes: detecting the load information of each edge cluster, the load information including: CPU usage, memory usage and disk space usage, and evaluating the load of each edge cluster according to the CPU usage, the memory usage and the disk space usage; detecting a network delay of each edge cluster, where the network delay is used to determine a distance between any two edge clusters; Record the image list that has been cached by each edge cluster, and obtain the image status that has been cached by each edge cluster.
8. The method for implementing dynamic image caching based on incremental snapshots according to claim 4, characterized in that: The incremental snapshot includes: creating a snapshot layer for the image through the COW file system, the edge cluster exports the incremental snapshot image for transmission, and the designated edge cluster receives and stores the incremental snapshot image.
9. The method for implementing dynamic image caching based on incremental snapshots according to claim 2, characterized in that: Determining whether caching is required in each of the cycles includes: Collect historical data on the cache status and image usage of each edge cluster and convert it into a time series. The preset Ci,j,t represents the cache status of image i on edge cluster j at time t; The preset Ui,j,t represents the usage of image i on edge cluster j at time t; the image cache status on each edge cluster is taken as a time series Sj,t Where nj is the number of mirrors in edge cluster j, which is n; Each of the time series is modeled by the ARIMA model to predict future time series values.
10. A system for implementing dynamic image caching based on incremental snapshots, applied to the method for implementing dynamic image caching based on incremental snapshots as claimed in any one of claims 1 to 9, characterized in that: include: A central cloud and several edge clusters, wherein the central cloud is deployed with an image-cloud module and a central storage module, and each edge cluster is deployed with an image-edge module and an edge storage module; The image-cloud module is used to: (1) Establishing grpc streaming bidirectional communication with the image-edge module; (2) receiving information reported by the image-edge module for policy judgment; (3) Recording image cache information for model generation; (4) judging whether the edge cluster needs to cache an image according to the model; (5) Determining the image cache mode according to the image cache information; (6) Support sending images to the image-edge module; The image-edge module is used to: (1) regularly updating the information of the edge cluster from the image-cloud module; (2) reporting the load information of the edge cluster to the image-cloud module of the other edge clusters; (3) used to receive or send images from other image-edge modules; used to receive images from the image-cloud module; (4) Use network measurement tools to detect the network delay from the current cluster to other clusters; (5) collecting load information of the current edge cluster; Wherein, if a new edge cluster is added, the central cloud updates the new edge cluster topology information to each edge cluster; The image-edge module of the central cloud regularly measures and records the network status of the edge cluster; The image-edge module periodically collects the edge cluster information, including CPU usage, memory usage, and disk space usage. After the information is collected, the network status and the cache status are reported to the image-cloud module; The image-cloud module collects information of each edge cluster and stores it for analysis.
Citation Information
Patent Citations
Mirror image pulling method and related product
CN115380269A
Mirror image warehouse distributed caching method and device
CN115687420A
Edge computing storage service providing method and device, electronic equipment and medium
CN115794319A
Mirror image caching method for edge computing virtual machine
CN116339914A
Cited By
Mirror image management method suitable for multi-cluster idle computing power scheduling
CN122111566A