Mirror image distribution method and device, equipment and storage medium
By obtaining multi-dimensional load data in a peer-to-peer network to calculate node weights, performing image sharding and parallel downloading of target sub-shards, and performing hash verification and reassembly, the problem of low transmission efficiency and centralization risk in existing Docker image distribution methods is solved, achieving efficient and stable image distribution.
Patent Information
- Application Number
- CN202511170045.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-20
- Publication Date
- 2025-11-11
AI Technical Summary
Existing Docker image distribution methods suffer from low transmission efficiency due to their static strategy and use of entire layers as transmission units. This makes them unable to meet the high-efficiency and fast image distribution requirements of large-scale containerized deployments. Furthermore, centralized designs pose single-point failure risks and bandwidth bottlenecks.
By acquiring multi-dimensional load data of each node in the peer-to-peer network, calculating node weights, performing mirror sharding and parallel downloading of target sub-shards, and performing hash verification and reassembly, the node selection and data transmission units are optimized.
It improves image download speed, avoids node overload, enhances system stability and efficiency, prevents data transmission errors or tampering, and optimizes transmission efficiency.
Smart Images

Figure CN120935197A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of cloud computing technology, and in particular to image distribution methods, apparatus, devices and storage media. Background Technology
[0002] In the current era of booming cloud computing and containerization technologies, Docker image distribution, as a crucial link in container deployment, directly impacts the performance and availability of the entire system. In the traditional model, Docker image distribution heavily relies on a centralized Registry server. While this architecture initially facilitated unified image management and distribution, its drawbacks have become increasingly apparent as container scale expands and application scenarios become more complex. The centralized design introduces a single point of failure risk into the entire distribution system. If the Registry server fails or is attacked, the image distribution service will be interrupted, affecting all services that rely on this service for container deployment. Simultaneously, image requests from all nodes converge on the central server, placing immense bandwidth pressure on it and creating a bandwidth bottleneck. This is especially severe in large-scale distributed systems where a large number of nodes simultaneously request images, leading to high image download latency and significantly impacting container deployment efficiency.
[0003] To alleviate the pressure of centralization, existing P2P (Peer-to-Peer) solutions have emerged, distributing the load on the central server through direct data transmission between nodes. However, these solutions have significant shortcomings in node selection strategies, often employing static approaches that cannot be adjusted in real-time according to dynamically changing network environments and the resource status of heterogeneous nodes. This results in some nodes becoming bottlenecks in data transmission due to poor network conditions or insufficient resources, impacting overall distribution efficiency. Furthermore, existing P2P solutions are not well-designed in terms of data transmission units, transmitting data in whole layers. For large layers like base images, the large data volume leads to slow downloads and fails to fully utilize the advantages of fragmented parallel transmission, thus failing to effectively improve image download speeds and making it difficult to meet the demands of modern large-scale containerized deployments for efficient and rapid image distribution. Summary of the Invention
[0004] This application provides a method, apparatus, device, and storage medium for distributing images, in order to solve the problem of low transmission efficiency caused by the adoption of static strategies and transmission in whole-layer units in the existing image distribution methods.
[0005] According to one aspect of the embodiments of this application, this application provides a method for distributing an image, the method comprising: acquiring multi-dimensional load data of each node in a peer-to-peer network, determining the node weight of the corresponding node based on the multi-dimensional load data of the node; sharding the image to obtain multiple sub-shards, storing each sub-shard in each node of the peer-to-peer network; when receiving an image download request, selecting multiple target nodes according to the node weights to download the target sub-shards in parallel; and verifying and reassembling each target sub-shard to obtain a target image corresponding to the image download request.
[0006] Optionally, before obtaining multi-dimensional load data of each node in the peer-to-peer network and determining the node weight of the corresponding node based on the multi-dimensional load data of the node, the method further includes: configuring the address of the scheduling center to each node; obtaining the registration data reported by each node through the scheduling center; and adding each node to the peer-to-peer network based on the registration data.
[0007] Optionally, the step of acquiring multi-dimensional load data of each node in the peer-to-peer network and determining the node weight of the corresponding node based on the multi-dimensional load data of the node includes: receiving heartbeat signals sent by each node in the peer-to-peer network, and acquiring the multi-dimensional load data sent by the node that sent the heartbeat signal, wherein the load data includes at least one of CPU data, bandwidth data, cache hit rate data, and network jitter data; and calculating the node weight of the corresponding node based on at least one of the CPU data, the bandwidth data, the cache hit rate data, and the network jitter data.
[0008] Optionally, the step of sharding the image to obtain multiple sub-shards and storing each sub-shard in each node of the peer-to-peer network includes: sharding each image layer in the constructed image based on a preset shard size to obtain multiple sub-shards for each image layer; and distributing each sub-shard to each node in the peer-to-peer network, including storing the same sub-shard on different nodes, wherein the scheduling center stores node distribution information for each sub-shard.
[0009] Optionally, when a mirror download request is received, selecting multiple target nodes to download target sub-shards in parallel based on the node weights includes: receiving the mirror download request initiated by the user; selecting the top N nodes with the highest node weights as the target nodes, where N is an integer greater than 1, and different target nodes store different sub-shards of each of the mirror layers; and downloading the target sub-shards of each mirror layer in parallel from each target node based on the mirror download request.
[0010] Optionally, the step of verifying and reassembling each target sub-segment to obtain the target image corresponding to the image download request includes: calculating the target sub-segment hash value of each target sub-segment; comparing the target sub-segment hash value of each target sub-segment with the sub-segment hash value of each target sub-segment stored in the scheduling center; if they are consistent, the hash verification of each target sub-segment passes; reassembling each target sub-segment after passing the hash verification layer by layer to obtain each reassembled image layer; performing layer consistency verification on each reassembled image layer; if the layer consistency verification passes, the target image corresponding to the image download request is obtained.
[0011] Optionally, after verifying and reassembling each of the target sub-segments to obtain the target image corresponding to the image download request, the method further includes: obtaining historical access data of the image from the user terminal within a preset time period; determining the highest frequency access period and the lowest frequency access period of the image from the user terminal based on the historical access data; distributing each of the sub-segments of the image to the edge nodes of the user terminal during the lowest frequency access period, so that the user terminal can pull each of the sub-segments of the image from the edge nodes during the highest frequency access period.
[0012] According to another aspect of the embodiments of this application, this application provides an image distribution apparatus, the apparatus comprising: a data acquisition module, configured to acquire multi-dimensional load data of each node in a peer-to-peer network, and determine the node weight of the corresponding node based on the multi-dimensional load data of the node; an image sharding module, configured to shard the image to obtain multiple sub-shards, and store each sub-shard in each node of the peer-to-peer network; a node selection module, configured to select multiple target nodes according to the node weights to download the target sub-shards in parallel when an image download request is received; and a shard reassembly module, configured to verify and reassemble based on each target sub-shard to obtain the target image corresponding to the image download request.
[0013] According to another aspect of the embodiments of this application, this application provides a computer device, including: a processor, a memory, and a network interface. The memory stores machine-readable instructions executable by the processor. When the computer device is running, the processor communicates with the memory through the network interface, and the processor executes the machine-readable instructions to perform the steps of the image distribution method as described above.
[0014] According to another aspect of the embodiments of this application, this application provides a computer-readable medium having processor-executable non-volatile program code that causes the processor to perform the steps of the image distribution method.
[0015] Compared with related technologies, the technical solutions provided in this application have the following advantages:
[0016] This application provides a method for image distribution. By acquiring multi-dimensional load data of each node in a peer-to-peer network, and combining this data to calculate the node weight of each node, the higher the node weight, the lower the current load of the node. This facilitates the selection of nodes with higher node weights as the nodes for image retrieval, thereby improving image download speed, avoiding node overload, and ultimately improving download efficiency. Before image download, the image is fragmented based on the peer-to-peer network, and each sub-fragment of the image is stored in each node of the peer-to-peer network. Then, the target sub-fragments are retrieved from nodes with higher node weights. This breaks through the limitation of traditional peer-to-peer networks using the entire layer as the transmission unit. Parallel retrieval of different target sub-fragments by multiple nodes is more conducive to accelerating download speed. In addition, the retrieval of each target sub-fragment is verified and reassembled to prevent data transmission errors or data tampering. Attached Figure Description
[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0018] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, those skilled in the art can obtain other drawings based on these drawings without creative effort.
[0019] Figure 1 This is a schematic diagram of the hardware environment for the image distribution method provided according to an embodiment of this application;
[0020] Figure 2 This is a flowchart illustrating an optional image distribution method provided according to an embodiment of this application;
[0021] Figure 3 This is an architecture diagram of an optional image provided according to an embodiment of this application;
[0022] Figure 4 This is an interaction diagram of an optional P2P-based Docker image distribution system provided according to an embodiment of this application;
[0023] Figure 5 This is a schematic diagram of an optional image distribution device according to an embodiment of this application;
[0024] Figure 6 This is a schematic diagram of an optional computer device structure provided for an embodiment of this application. Detailed Implementation
[0025] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0026] To address the problems mentioned in the background art, according to one aspect of the embodiments of this application, an embodiment of a method for distributing an image is provided.
[0027] like Figure 1 As shown, the above image distribution method can be applied to, for example... Figure 1 The hardware environment shown is described. The system architecture 100 of the hardware environment includes a terminal device 101 and a server 103. The server 103 is connected to the terminal 101 via a network and can be used to provide services to the terminal or clients installed on the terminal. A database 105 can be set up on the server or independently of the server to provide data storage services to the server 103. The network can include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0028] Users can use terminal device 101 to interact with server 103 via a network to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as web browser applications, search applications, instant messaging tools, etc. Terminal device 101 can be various electronic devices with a display screen that support web browsing, including but not limited to smartphones, tablets, laptops, laptops, and desktop computers. Server 103 can be a server that provides various services, such as a backend server that supports the pages displayed on terminal device 101.
[0029] It should be noted that the image distribution method provided in this application embodiment is generally executed by a server and / or a terminal device, and correspondingly, the image distribution device is generally set in the server and / or the terminal device.
[0030] like Figure 2 As shown, Figure 2 A flowchart illustrating an image distribution method provided in this embodiment of the invention. Taking the image distribution method being executed by a server as an example, the image distribution method includes the following steps:
[0031] Step S202: Obtain multi-dimensional load data of each node in the peer-to-peer network, and determine the node weight of the corresponding node based on the multi-dimensional load data of the node.
[0032] Peer-to-peer (P2P) networks are a decentralized distributed network architecture. Their core principle is that all nodes in the network have equal status; they are both resource requesters and resource providers. That is, a node can be both a user and a server. They feature no single point of failure, high network stability, strong resistance to attacks, and wide resource distribution, avoiding server bottlenecks and improving overall efficiency. In a P2P network, each node can be a server host or virtual machine, including user terminal devices, as well as other terminal devices joining the P2P network, and a management node that schedules data between the terminal devices and user terminals. The management node can refer to a scheduling center.
[0033] Load data includes, but is not limited to, node CPU data, bandwidth data, cache hit rate data, storage data, and network jitter data, such as CPU utilization and idle bandwidth. Load data for each node can be dynamically retrieved in real time. After obtaining multi-dimensional load data for each node in the P2P network, the node weight of each node can be dynamically calculated based on its load data, thus calculating the node weight of all nodes. The maximum node weight is 1. A higher node weight indicates a lower current load, and a lower node weight is beneficial for faster image retrieval and improved image download speed.
[0034] Step S204: The image is split into multiple sub-shards, and each sub-shard is stored in each node of the peer network.
[0035] Here, "image" can refer to a Docker image, which can be composed of multiple read-only layers stacked together, such as Docker's UnionFS. In an image, a read-only layer can also refer to an image layer. Each image layer represents a file system change, such as installing dependencies or copying files.
[0036] Combination Figure 3 As shown, Figure 3 This is a schematic diagram of the overall structure of an optional Docker image provided in this embodiment. Figure 3In a Docker image, the outermost layer represents the image boundary. Internally, it consists of multiple stacked image layers, each representing a component of the image. Layers 1 and 2 are located at the bottom of the image and can contain the basic operating system files and configurations. For example, Layer 1 might be a base layer based on a Linux distribution, containing basic system files, libraries, and commands. Layer 2 adds additional system configurations or software packages on top of Layer 1, further customizing the operating system environment. Segment 5 and Segment 6 can be two sub-shards obtained through sharding; for example, layer 3 in native Docker can be sharded into Segment 5 and Segment 6.
[0037] In this embodiment, to accelerate the download speed of the image, each image layer can be fragmented, resulting in multiple sub-fragments for each image layer. For example, if the image contains 3 image layers, each image layer can be divided into 2 sub-fragments, resulting in a total of 6 sub-fragments. Of course, this is just an example; each image layer can also be divided into more sub-fragments, such as 50, 100, etc. The resulting sub-fragments are distributed and stored across nodes in a P2P network. Identical sub-fragments can be stored on multiple nodes; for example, nodes 1, 2, and 3 all store sub-fragment 'a'.
[0038] Step S206: When a mirror download request is received, multiple target nodes are selected according to the node weights to download the target sub-shards in parallel.
[0039] The user client can send an image download request to the scheduling center to request the retrieval of sub-fragments from various nodes in the P2P network, and then reassemble the retrieved sub-fragments to obtain the desired image for download. Upon receiving the image download request, the scheduling center dynamically retrieves the load data of each node in the current P2P network to calculate node weights, prioritizing nodes with higher weights as target nodes. The center can return a list of locked target nodes to the user client, where some or all of the image's sub-fragments are already cached.
[0040] Furthermore, the user can send segmented download requests to each of the returned target nodes. Each target node may store different sub-segments of the image, and the user can download different target sub-segments from multiple target nodes in parallel, which can improve the download speed.
[0041] S208, based on each of the target sub-segments, verification and reassembly are performed to obtain the target image corresponding to the image download request.
[0042] In order to ensure the security and accuracy of the data and prevent transmission errors or tampering after downloading multiple target sub-fragments, each target sub-fragment can be verified. The verification includes dual verification formed by sub-fragment verification and hierarchical verification.
[0043] In this embodiment, sub-shard verification can be performed by comparing the downloaded target sub-shard with the sub-shards that are identical to the target sub-shard stored in the scheduling center. This includes verification by comparing the hash values of the two sub-shards. For example, the hash value of the target sub-shard a is calculated and the hash value of the sub-shard a′ is calculated. If they are identical, the verification passes; otherwise, the verification fails. Here, the target sub-shard a and the sub-shard a′ are actually the same sub-shard stored in different nodes.
[0044] Furthermore, layer verification refers to verifying the consistency of the image layers composed of target sub-shards. It's important to note that layer reassembly and layer verification are native features of Docker images. Once each target sub-shard passes verification, they are reassembled into layers. The integrity of the layer is verified by calculating its hash value and comparing it to the hash value recorded in the image manifest. If multiple reassembled layers pass verification, they are reassembled into a complete target image, thus completing the download.
[0045] In this embodiment of the invention, by acquiring multi-dimensional load data of each node in the peer-to-peer network, and combining the multi-dimensional load data to calculate the node weight of each node, the higher the node weight, the lower the current load of the node. This is beneficial for selecting nodes with higher node weights as the nodes for image retrieval, thereby improving the image download speed, avoiding node overload, and thus improving download efficiency. Before image download, image fragmentation is performed based on the peer-to-peer network, and each sub-fragment of the image is stored in each node of the peer-to-peer network. Then, the target sub-fragments are retrieved from the nodes with higher node weights. This breaks through the limitation of traditional peer-to-peer networks using the entire layer as the transmission unit. Pulling different target sub-fragments in parallel by multiple nodes is more conducive to accelerating the download speed. In addition, the verification and reassembly of each retrieved target sub-fragments can prevent data transmission errors or data tampering.
[0046] In some optional embodiments, prior to step S202 above, the method further includes:
[0047] Configure the address of the scheduling center on each node;
[0048] The scheduling center obtains the registration data reported by each node and adds each node to the peer-to-peer network based on the registration data.
[0049] In this embodiment, in order to enable each node to join the P2P network, the address of the scheduling center can be pre-configured for each node. After each node configures the address of the scheduling center, it can establish communication with the scheduling center. Then, each node can report its registration data to the scheduling center, and the scheduling center adds each node to the P2P network. The registration data may include registration identifier, IP address, node type, network bandwidth, port number, etc.
[0050] In this embodiment, by configuring the scheduling center address for each node, a connection channel can be established between the node and the scheduling center. Then, the registration data reported by each node can be obtained through the scheduling center, and each node can be accurately added to the P2P network based on this data. This enables the efficient and orderly construction of a complete P2P network architecture, ensuring that each node can smoothly integrate into the network and achieve mutual communication and cooperation.
[0051] In some optional embodiments, step S202 above includes:
[0052] Receive heartbeat signals sent by each node in the peer-to-peer network, and acquire multi-dimensional load data sent by the node that sent the heartbeat signal. The load data includes at least one of CPU data, bandwidth data, cache hit rate data, and network jitter data.
[0053] The node weight of the corresponding node is calculated based on at least one of the CPU data, the bandwidth data, the cache hit rate data, and the network jitter data.
[0054] Combination Figure 4 As shown, Figure 4 This embodiment provides an interactive diagram of an optional P2P-based Docker image distribution system. Figure 4 This involves the interaction between the user, the scheduling center, and the nodes (Peer1, Peer2, Peer3) in the P2P network. Through a heartbeat mechanism, nodes Peer1, Peer2, and Peer3 periodically send heartbeat signals to the scheduling center to indicate their activity and participation in image distribution and load balancing. This helps the scheduling center maintain the availability of nodes Peer1, Peer2, and Peer3. Conversely, if a node fails to send a heartbeat signal to the scheduling center within a preset time period, it can be preliminarily considered unavailable, potentially due to node failure or overload. This prevents such nodes from participating in image distribution and load balancing, reducing the impact on image distribution efficiency.
[0055] Furthermore, the scheduling center receives multi-dimensional load data sent by the nodes that send heartbeat signals. This load data includes at least one of the following: CPU data, bandwidth data, cache hit rate data, and network jitter data. It may also include, but is not limited to, storage data, memory data, thread data, and task execution data. For example, storage data may include disk space utilization and disk I / O load; memory data may include memory utilization; thread data may include the running status of key processes, the number of threads, and their load; and task execution data may include the task queue length and task execution success rate.
[0056] In some possible embodiments, load data includes CPU data, bandwidth data, cache hit rate data, and network jitter data. When calculating the node weight based on CPU data, bandwidth data, cache hit rate data, and network jitter data, CPU data can refer to CPU load data, and bandwidth data refers to idle bandwidth data. The calculation formula is: Node weight = a*(1 / CPU load data) + b*(idle bandwidth data) + c*(cache hit rate data) + d*(network jitter data), where load data is expandable, and a, b, c, and d are weight coefficients, the sum of which is 1. The higher the node weight, the more likely it is to be selected as the target node.
[0057] In this embodiment, by combining the load data of nodes to dynamically calculate nodes, and taking into account dynamic indicators such as CPU load, idle bandwidth, network latency, and cache hit rate of nodes, reinforcement learning is used to optimize node weights. This enables a more comprehensive analysis and accurate prioritization of nodes with high node weights as target nodes for sub-sharding, thereby accelerating the retrieval speed, improving image download efficiency, and avoiding faulty or overloaded nodes in real time.
[0058] In some optional embodiments, step S204 above includes:
[0059] The image layers in the constructed image are split based on the preset split size to obtain multiple sub-shards of each image layer;
[0060] Distributing the sub-shards to nodes in the peer-to-peer network includes storing the same sub-shard on different nodes, and the scheduling center stores the node distribution information of each sub-shard.
[0061] In this embodiment, the shard size can be predetermined. In the storage and transmission of Docker images, the shard size can be 4MB, 8MB, etc. Docker images adopt a layered image structure and content-addressable storage method. A smaller shard size helps to improve the cache utilization and transmission efficiency of images, especially when pulling and pushing images in a distributed environment.
[0062] Furthermore, the built Docker image is sharded according to the shard size, and each image layer within the Docker image can be sharded separately. This results in multiple sub-shards for each layer of the Docker image. For example, image layer Layer1 can be divided into 3 sub-shards based on a size of 4MB. Further, each sub-shard is stored across nodes in a P2P network, and the scheduling center stores the nodes containing each sub-shard. This information is used to inform the user of the node containing the desired target sub-shard when an image download request is received. The same sub-shard can be distributed across multiple nodes; for example, sub-shard 'a' can be stored on nodes peer1, peer2, and peer3. Distributed storage across different nodes allows for the selection of faster and more efficient target nodes for downloading sub-shards based on the load data of each node. Of course, distributing sub-shards to nodes also includes storing each sub-shard in only one node. For example, storing sub-shard a in node Peer1, sub-shard b in node Peer2, and so on, in a cyclical manner until all sub-shards are stored.
[0063] In this embodiment, by sharding each image layer and distributing the sub-shards across nodes in the P2P network, data reliability and availability are improved, avoiding the adverse effects of single node failure or data corruption. Furthermore, storage scalability is enhanced to meet the demands of massive image storage, allowing for the addition or removal of nodes as needed. In addition, further splitting each layer of the Docker image into fixed-size shards overcomes the limitations of traditional P2P which uses entire layers as transmission units, optimizing transmission efficiency. By enabling multiple nodes to download sub-shards in parallel, network bandwidth is fully utilized, accelerating image acquisition.
[0064] In some optional embodiments, step S206 specifically includes:
[0065] Receive the image download request initiated by the user client;
[0066] The top N nodes with the highest node weights are selected as the target nodes, where N is an integer greater than 1. Different target nodes store different sub-shards of each of the mirror layers.
[0067] Based on the image download request, the target sub-fragments of each image layer are downloaded in parallel from each of the target nodes.
[0068] Among them, combined Figure 4As shown, the client, acting as the initiator of the image download, can send an image download request to the scheduling center to obtain the required Docker image. Upon receiving the request, the scheduling center selects the top N nodes with the highest node weights as target nodes for sub-shards based on real-time calculated node weights, and returns this list to the client. For example, it might select the top 3 nodes with the highest node weights as target nodes and return them to the client. Each target node returned by the scheduling center stores different target sub-shards for each image layer of the Docker image the client needs to download; that is, the sum of the target sub-shards stored in each target node represents the Docker image the client needs to download.
[0069] Furthermore, after receiving the target node returned by the scheduling center, the user terminal can send segmented download requests to each target node to request the download of different target sub-segments. For example, it can send Downloadsegment1 to target node peer1, Downloadsegment2 to target node peer2, and Downloadsegment3 to target node peer3. Each node stores different sub-segments of the image. The user terminal can download different target sub-segments from multiple nodes in parallel, thereby improving the download speed. Then, the sub-segments downloaded from different nodes are reassembled to obtain a reconstructed image, completing the download.
[0070] In some examples, load data can be prioritized. Higher priority data is considered more frequently as a target node. For instance, idle bandwidth data has a higher priority than CPU utilization data. If node peer1 and node peer2 have the same node weight, the node with higher priority and larger idle bandwidth data can be selected as the target node. This facilitates the rapid selection of the target node.
[0071] In other examples, when different nodes store the same target sub-shard, one of them can be selected as the target node by combining the node weights. For example, if nodes peer1 and peer2 both contain target sub-shard a, and node peer1 is currently idle while node peer2 has high CPU utilization, then node peer1 has a higher node weight than node peer2. Therefore, node peer1 will be selected as the target node, and target sub-shard a will be downloaded from node peer1 first.
[0072] In this embodiment, the node weight of each node is calculated by combining multi-dimensional load data. The higher the node weight, the lower the current load of the node. This is beneficial for selecting the target node with a higher node weight as the node for pulling the sub-shards of the image. By downloading the target sub-shards in parallel by multiple target nodes, the download speed of the image is improved, node overload is avoided, and thus the download efficiency is improved.
[0073] In some optional embodiments, step S208 above includes:
[0074] Calculate the target fragment hash value for each of the target sub-fragments;
[0075] The target sub-segment hash value of each target sub-segment is compared with the segment hash value of each target sub-segment stored in the scheduling center. If they are consistent, the hash verification of each target sub-segment passes.
[0076] After the hash verification is passed, each target sub-segment is reassembled layer by layer to obtain each reassembled mirror layer;
[0077] Perform layer consistency verification on each of the reconstructed image layers. If the layer consistency verification passes, the target image corresponding to the image download request is obtained.
[0078] In this embodiment, for verification, the target sub-segment can be hashed to calculate its target sub-segment hash value. Then, the target sub-segment hash value of each target sub-segment is compared with the sub-segment hash value stored in the scheduling center. If they match, the hash verification of the target sub-segment passes. Through the above hash verification, the verification of each target sub-segment can be achieved until all target sub-segment hash verifications pass. If some target sub-segment hash verifications fail, a fragment verification failure report can be sent to the scheduling center, and further hierarchical verification will not proceed. For example, if the target sub-segment a obtains a sub-segment hash value of abc123... from the scheduling center, and the target sub-segment hash value calculated by the user is also abc123..., then these two hash values match, indicating that target sub-segment a did not experience data corruption or loss during the download process and is a complete and correct fragment. For example, the target sub-segment b obtains the segment hash value def456... from the scheduling center, while the target segment hash value calculated by the user is xyz789... The two hash values are inconsistent, which indicates that the target sub-segment b may have encountered a problem during the download process, causing the file content to change and the verification to fail.
[0079] In some examples, after the hash verification of each target sub-shard to be downloaded by the user passes, layer reassembly can be performed to obtain the reassembled Docker image layers. Further, layer consistency verification can be performed on each reassembled image layer; layer verification is a native Docker function and will not be elaborated upon here. After all layer consistency verifications pass, the images are reassembled into a complete image, yielding the target image for the corresponding image download request.
[0080] In this embodiment, based on layered fragment reassembly, combined with dual verification of fragment-level hash verification and Docker layer verification, it is possible to accurately locate and quickly repair damaged data blocks in transmission or storage at the fragment level, ensuring that the image fragments are complete and error-free. It can also verify the integrity and consistency of each layer from the perspective of the Docker image layer structure, preventing data errors between layers during image building or modification, and comprehensively ensuring the quality and reliability of the image.
[0081] In some optional embodiments, after step S208 above, the method further includes:
[0082] Obtain historical access data of the mirror from the user terminal within a preset time period;
[0083] Based on the historical access data, determine the highest and lowest frequency access periods for the user's access to the mirror;
[0084] During the lowest frequency access period, each of the sub-shards of the image is distributed to the edge nodes of the user terminal, so that the user terminal can pull each of the sub-shards of the image from the edge nodes during the highest frequency access period.
[0085] In this embodiment, the edge node can be a lightweight device close to the data source or end user, and in this embodiment, it can refer to a specific node within a region. To achieve advance deployment and accelerate the download efficiency of the user terminal, historical access data of the user terminal to the image within a preset time period can be obtained in advance. For example, historical access data of the user terminal to the image for each hour in the past 24 hours can be obtained. The historical access data can include access frequency, access duration, image download time, etc.
[0086] Furthermore, the highest and lowest frequency access periods for the image by the user can be determined based on the access frequency. Then, during the lowest frequency access period, each sub-shard of the image is distributed to the edge nodes of the user, so that the user can pull each sub-shard of the image from the edge nodes during the highest frequency access period. This enables fast image download through pre-deployment. For example, in scenario A, historical access data shows that from 8:00 AM to 10:00 AM the previous day, the camera AI processing container located in the industrial park would batch pull the image tensorflow-lite:2.8-arm32v7 for real-time image recognition. This image contains 3 layers, each layer is sharded into 10 sub-shards of 4MB each, for a total of 30 sub-shards. By analyzing historical logs, the scheduling center found that this image experiences the highest frequency of access during the morning peak hours (accounting for 80% of all requests) in the industrial park, and the least access and lowest load occurs at 3:00 AM (the lowest frequency of access). Therefore, at 3:00 AM, the scheduling center proactively transmits the 30 sub-shards to the edge node x in the industrial park. When the camera container starts at 8:00 AM, it directly pulls the sub-shards through node x. By using cross-public network to local P2P, the download speed can be reduced from the original 2 minutes to 5 seconds, greatly shortening the download time, speeding up the download, and improving download efficiency.
[0087] In this embodiment, through intelligent preheating and cache optimization, the high-frequency access period of the image is predicted based on historical access data. The sub-fragments of the image are pushed to the edge nodes of the client in advance during the lowest frequency access period, so that the client can pull the sub-fragments of the image from the edge nodes during the highest frequency access period, which greatly shortens the download time, speeds up the download, and improves the download efficiency.
[0088] According to another aspect of the embodiments of this application, such as Figure 5 As shown, corresponding to the image distribution method in the above embodiments, this embodiment provides an image distribution device, the device comprising:
[0089] The data acquisition module 501 is used to acquire multi-dimensional load data of each node in the peer-to-peer network and determine the node weight of the corresponding node based on the multi-dimensional load data of the node.
[0090] The image sharding module 503 is used to shard the image to obtain multiple sub-shards, and store each of the sub-shards in each node of the peer network;
[0091] The node selection module 505 is used to select multiple target nodes in parallel to download the target sub-segment in accordance with the node weights when a mirror download request is received.
[0092] The fragment reassembly module 507 is used to verify and reassemble each of the target sub-fragments to obtain the target image corresponding to the image download request.
[0093] It should be noted that in this embodiment, the data acquisition module 501 can be used to execute step S202 in this application embodiment, the mirror sharding module 503 in this embodiment can be used to execute step S204 in this application embodiment, the node selection module 505 in this embodiment can be used to execute step S206 in this application embodiment, and the sharding reassembly module 507 in this embodiment can be used to execute step S208 in this application embodiment.
[0094] In some alternative embodiments, the apparatus further includes: an address configuration module for configuring the address of the scheduling center to each node; and a node joining module for obtaining registration data reported by each node through the scheduling center and joining each node to the peer-to-peer network based on the registration data.
[0095] In some optional embodiments, the data acquisition module 501 includes: a load data acquisition submodule, configured to receive heartbeat signals sent by each node in the peer-to-peer network and acquire multi-dimensional load data sent by the node sending the heartbeat signal, wherein the load data includes at least one of CPU data, bandwidth data, cache hit rate data, and network jitter data; and a first calculation submodule, configured to calculate the node weight of the corresponding node based on at least one of the CPU data, the bandwidth data, the cache hit rate data, and the network jitter data.
[0096] In some optional embodiments, the image sharding module 503 includes: a sharding submodule, used to shard each image layer in the constructed image based on a preset sharding size to obtain multiple sub-shards for each image layer; and a storage submodule, used to distribute and store each sub-shard to each node in the peer-to-peer network, including storing the same sub-shard on different nodes, wherein the scheduling center stores node distribution information of each sub-shard.
[0097] In some optional embodiments, the node selection module 505 includes: a receiving submodule, configured to receive the image download request initiated by the user terminal; select the top N nodes with the highest node weights as the target nodes, where N is an integer greater than 1, and different target nodes store different sub-shards of each image layer; and a download submodule, configured to download the target sub-shards of each image layer in parallel from each target node based on the image download request.
[0098] In some optional embodiments, the fragment reassembly module 507 includes: a second calculation submodule, used to calculate the target fragment hash value of each target sub-fragment; a comparison submodule, used to compare the target fragment hash value of each target sub-fragment with the fragment hash value of each target sub-fragment stored in the scheduling center, and if they are consistent, the hash verification of each target sub-fragment passes; a reassembly submodule, used to reassemble each target sub-fragment that has passed the hash verification by layer to obtain each reassembled image layer; and a verification submodule, used to perform layer consistency verification on each reassembled image layer, and if the layer consistency verification passes, the target image corresponding to the image download request is obtained.
[0099] In some optional embodiments, the apparatus further includes: a historical data acquisition module, configured to acquire historical access data of the user terminal to the mirror within a preset time period; a time period determination module, configured to determine the highest frequency access time period and the lowest frequency access time period of the user terminal to the mirror based on the historical access data; and a distribution module, configured to distribute each of the sub-shards of the mirror to the edge nodes of the user terminal during the lowest frequency access time period, so that the user terminal can pull each of the sub-shards of the mirror from the edge nodes during the highest frequency access time period.
[0100] It should be noted that the examples and application scenarios implemented by the above modules, sub-modules, and corresponding steps are the same, but are not limited to the content disclosed in the above embodiments. It should also be noted that the above modules and sub-modules, as part of the system, can run in the hardware environment of the image distribution device, and can be implemented in software or hardware.
[0101] According to another aspect of the embodiments of this application, this application provides a computer device, such as... Figure 6 As shown, it includes a memory 601, a processor 603, a communication interface 605, and a communication bus 607. The memory 601 stores a computer program that can run on the processor 603. The memory 601 and the processor 603 communicate through the communication interface 605 and the communication bus 607. When the processor 603 executes the computer program, it implements the steps of the above-mentioned image distribution method.
[0102] The memory and processor in the aforementioned computer equipment communicate with each other via a communication bus and a communication interface. The communication bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc.
[0103] The aforementioned memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.
[0104] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0105] According to another aspect of the embodiments of this application, a computer program product or computer program is also provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the steps of the image distribution method in any of the above embodiments.
[0106] Optionally, in this embodiment, the computer-readable medium is configured to store program code for the processor to execute the steps of the image distribution method described in the above embodiments, wherein the steps of the image distribution method specifically include:
[0107] S202. Obtain multi-dimensional load data of each node in the peer-to-peer network, and determine the node weight of the corresponding node based on the multi-dimensional load data of the node.
[0108] S204. The image is fragmented to obtain multiple sub-fragments, and each of the sub-fragments is stored in each node of the peer-to-peer network.
[0109] S206. When a mirror download request is received, multiple target nodes are selected according to the node weights to download the target sub-fragments in parallel.
[0110] S208. Verify and reassemble each of the target sub-segments to obtain the target image corresponding to the image download request.
[0111] Optionally, specific examples in this embodiment can refer to the examples described in the above embodiments, and will not be repeated here. Furthermore, in the specific implementation of this application embodiment, the above embodiments can be consulted, and corresponding technical effects can be achieved.
[0112] It is understood that the embodiments described herein can be implemented in hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described herein, or combinations thereof. For software implementation, the techniques described herein can be implemented by units that perform the functions described herein. Software code can be stored in memory and executed by a processor. The memory can be implemented in the processor or external to the processor.
[0113] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0114] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the division of modules is merely a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.
[0115] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, essentially or in other words, the parts that contribute to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.
[0116] It should be noted that, in this document, relational terms such as "first," "second," etc., are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprise," "include," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprises a…" does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0117] The above description is merely a specific embodiment of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A method for distributing images, characterized in that, The method includes: Obtain multi-dimensional load data of each node in the peer-to-peer network, and determine the node weight of the corresponding node based on the multi-dimensional load data of the node. The image is split into multiple sub-shards, and each sub-shard is stored in each node of the peer-to-peer network. When a mirror download request is received, multiple target nodes are selected according to the node weights to download the target sub-fragments in parallel; Based on the verification and reassembly of each target sub-segment, the target image corresponding to the image download request is obtained.
2. The image distribution method according to claim 1, characterized in that, Before acquiring multi-dimensional load data of each node in the peer-to-peer network and determining the node weight of the corresponding node based on the multi-dimensional load data, the method further includes: Configure the address of the scheduling center on each node; The scheduling center obtains the registration data reported by each node and adds each node to the peer-to-peer network based on the registration data.
3. The image distribution method according to claim 1, characterized in that, The step of acquiring multi-dimensional load data of each node in the peer-to-peer network and determining the node weight of the corresponding node based on the multi-dimensional load data includes: Receive heartbeat signals sent by each node in the peer-to-peer network, and acquire multi-dimensional load data sent by the node that sent the heartbeat signal. The load data includes at least one of CPU data, bandwidth data, cache hit rate data, and network jitter data. The node weight of the corresponding node is calculated based on at least one of the CPU data, the bandwidth data, the cache hit rate data, and the network jitter data.
4. The image distribution method according to claim 2, characterized in that, The process of splitting the image into multiple sub-shards and storing each sub-shard in each node of the peer-to-peer network includes: The image layers in the constructed image are split based on the preset split size to obtain multiple sub-shards of each image layer; Distributing the sub-shards to nodes in the peer-to-peer network includes storing the same sub-shard on different nodes, and the scheduling center stores the node distribution information of each sub-shard.
5. The image distribution method according to claim 4, characterized in that, When a mirror download request is received, selecting multiple target nodes based on the node weights to download the target sub-shards in parallel includes: Receive the image download request initiated by the user client; The top N nodes with the highest node weights are selected as the target nodes, where N is an integer greater than 1. Different target nodes store different sub-shards of each of the mirror layers. Based on the image download request, the target sub-fragments of each image layer are downloaded in parallel from each of the target nodes.
6. The image distribution method according to claim 1, characterized in that, The process of verifying and reassembling each of the target sub-segments to obtain the target image corresponding to the image download request includes: Calculate the target fragment hash value for each of the target sub-fragments; The target sub-segment hash value of each target sub-segment is compared with the segment hash value of each target sub-segment stored in the scheduling center. If they are consistent, the hash verification of each target sub-segment passes. After the hash verification is passed, each target sub-segment is reassembled layer by layer to obtain each reassembled mirror layer; Perform layer consistency verification on each of the reconstructed image layers. If the layer consistency verification passes, the target image corresponding to the image download request is obtained.
7. The method for distributing an image according to any one of claims 1 to 6, characterized in that, After verifying and reassembling each of the target sub-shards to obtain the target image corresponding to the image download request, the method further includes: Obtain historical access data of the mirror from the user terminal within a preset time period; Based on the historical access data, determine the highest and lowest frequency access periods for the user's access to the mirror; During the lowest frequency access period, each of the sub-shards of the image is distributed to the edge nodes of the user terminal, so that the user terminal can pull each of the sub-shards of the image from the edge nodes during the highest frequency access period.
8. A mirror distribution device, characterized in that, The device includes: The data acquisition module is used to acquire multi-dimensional load data of each node in the peer-to-peer network, and determine the node weight of the corresponding node based on the multi-dimensional load data of the node. The image sharding module is used to shard the image to obtain multiple sub-shards, and store each of the sub-shards in each node of the peer network; The node selection module is used to select multiple target nodes in parallel to download the target sub-segments in accordance with the node weights when a mirror download request is received. The fragment reassembly module is used to verify and reassemble each target sub-fragment to obtain the target image corresponding to the image download request.
9. A computer device, comprising: A processor, a memory, and a network interface, wherein the memory stores machine-readable instructions executable by the processor, characterized in that: when the computer device is running, the processor communicates with the memory via the network interface, and the processor executes the machine-readable instructions to perform the steps of the image distribution method as described in any one of claims 1 to 7.
10. A computer-readable medium having processor-executable non-volatile program code, characterized in that, The program code causes the processor to perform the steps of the image distribution method according to any one of claims 1 to 7.
Citation Information
Cited By
Large model file distribution acceleration method and system based on stateless peer-to-peer network
CN122027617A