Mirror data processing method, DPU and system
By processing mirror queries and block file queries in the target network of multiple DPU nodes, and using distributed storage and point-to-point network technology, the problems of bandwidth bottlenecks and network delays of centralized mirror warehouses are solved, and efficient mirror data download and load isolation are achieved.
Patent Information
- Application Number
- CN202410478153.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-19
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2044-04-19
AI Technical Summary
Existing centralized mirrored warehouses are prone to bandwidth bottlenecks in large-scale or high traffic conditions, resulting in reduced download and upload speeds, affecting delivery efficiency, and may be affected by network latency and server load.
By receiving and processing mirror query keywords and block file query keywords in a target network composed of multiple DPU nodes communicating with each other and sharing data, distributed storage and point-to-point network technology are used to realize distributed storage and efficient download of mirror data.
Effectively reduce the load burden of the host, improve the service quality of the application, realize load isolation, and improve the network transmission performance of mirrored data, thereby improving download efficiency and reliability.
Smart Images

Figure CN118283059B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to an image data processing method, DPU and system. Background Art
[0002] At present, mainstream image repositories often adopt a centralized architecture, which often has many limitations during use. For example, when users download or upload images, they need to rely on the bandwidth of the central warehouse. In large-scale or high-traffic situations, the bandwidth of the central server becomes a bottleneck, resulting in a decrease in download and upload speeds, affecting delivery efficiency. In addition, since users need to communicate with the central server or private warehouse through the Internet, they may be affected by network latency. This may cause problems in distributed teams, cross-geographic deployments, or in unstable network environments. In addition, the server will occupy input / output operations in the computer network during the image download process, affecting business applications running on the server.
[0003] Therefore, there is an urgent need to solve the problems of slow image download speed caused by existing centralized warehouses and the impact of the download process on business applications. Summary of the invention
[0004] In view of this, embodiments of the present application provide a mirror data processing method, a DPU, and a system to eliminate or improve one or more defects existing in the prior art.
[0005] One aspect of the present application provides a mirror data processing method, including:
[0006] In a target network composed of a plurality of DPU nodes that communicate with each other and share data, a keyword sent by any of the DPU nodes is received, wherein the keyword includes: an image query keyword and / or a block file query keyword; each block file corresponding to the image data is pre-distributed to each of the DPU nodes for distributed storage;
[0007] The received keyword is used as the current target keyword, and the block file information corresponding to the target keyword is searched from the correspondence between each keyword and each block file information pre-stored in the target network;
[0008] The storage location identifier pointed to by the block file information corresponding to the target keyword is returned to the DPU node that issued the target keyword, so that the DPU node downloads each block file corresponding to the target keyword from other DPU nodes corresponding to the storage location identifier and forms the mirror data.
[0009] In some embodiments of the present application, before receiving a keyword sent by any of the DPU nodes in the target network composed of multiple DPU nodes that communicate with each other and share data, the method further includes:
[0010] Receive mirror data;
[0011] Splitting the mirror data to obtain a plurality of block files, and setting a unique identifier for the mirror data and each of the block files;
[0012] Generate seed information corresponding to each of the block files, wherein the seed information includes a unique identifier of the block file, and also includes a name of the mirror data to which the block file belongs, a size of the mirror data, and a unique identifier of the mirror data;
[0013] Distributing each of the block files and the corresponding seed information to each of the DPU nodes for distributed storage;
[0014] And, using the unique identifier of each block file as a block file query keyword, and / or using the unique identifier or name of the mirror data as the mirror query keyword, a one-to-one correspondence between the unique identifier corresponding to each block file, the name of the mirror data to which it belongs, the size of the mirror data and the unique identifier of the mirror data is stored.
[0015] In some embodiments of the present application, before receiving the mirrored data, the method further includes:
[0016] Get the current network topology type of the target network;
[0017] Determine the current propagation path corresponding to the target network according to the network topology type;
[0018] The propagation path is distributed to each of the DPU nodes in the target network, so that the DPU node that issues the target keyword searches for other DPU nodes corresponding to the storage location identifier in the target network based on the propagation path when or after receiving the storage location identifier, and downloads each block file corresponding to the target keyword from the other DPU nodes corresponding to the storage location identifier.
[0019] In some embodiments of the present application, it also includes:
[0020] If the propagation path shows that there is a DPU node whose adjacent DPU node is not unique, then the DPU node is used as the current target node;
[0021] Acquire current status information of multiple adjacent DPU nodes corresponding to the target node, and based on the current status information of each of the adjacent DPU nodes;
[0022] Selecting one of the adjacent DPU nodes corresponding to the target node as a current neighbor node of the target node;
[0023] The selection information corresponding to the neighbor node is sent to the target node, so that the target node selects the neighbor node from the plurality of adjacent DPU nodes.
[0024] In some embodiments of the present application, the state information includes: at least one of the current bandwidth, load information, location information, availability information, and popularity information of the neighboring DPU node;
[0025] The availability information includes: the number and type of block files currently stored by the adjacent DPU node, and the download volume for the stored block files;
[0026] The popularity information includes: the download frequency or the sharing number of the block file currently stored in the adjacent DPU node in the target network.
[0027] In some embodiments of the present application, the DPU node that issues the target keyword verifies the integrity and correctness of each block file corresponding to the target keyword based on preset verification rules when or after downloading the block files, and combines the block files after the verification passes to obtain the mirror data.
[0028] In some embodiments of the present application, the network topology type includes: a tree network, a mesh network, or a point-to-point network;
[0029] Correspondingly, if the current network topology type of the target network is the tree network, the propagation path corresponding to the tree network includes: propagating from top to bottom in the tree structure formed by each DPU node in the tree network;
[0030] If the current network topology type of the target network is the mesh network, the propagation path corresponding to the mesh network includes: propagation through each adjacent DPU node in the mesh network;
[0031] If the current network topology type of the target network is the point-to-point network, the propagation path corresponding to the mesh network includes: mutual propagation between the DPU nodes that communicate with each other in the point-to-point network.
[0032] A second aspect of the present application provides a mirror data processing method, comprising:
[0033] In a target network composed of a plurality of DPU nodes that communicate with each other and share data, a download request for image data is received, and a corresponding keyword is obtained from the download request, the keyword including: an image query keyword and / or a block file query keyword; each block file corresponding to the image data is pre-distributed to each of the DPU nodes for distributed storage;
[0034] Sending the keyword to the mirror service node corresponding to the target network, so that the mirror service node uses the received keyword as the current target keyword, searches for the block file information corresponding to the target keyword from the correspondence between each keyword and each block file information pre-stored in the target network, and then returns the storage location identifier pointed to by the block file information corresponding to the target keyword;
[0035] The storage location identifier is received, and each block file corresponding to the target keyword is downloaded from other DPU nodes corresponding to the storage location identifier to form the mirror data.
[0036] A third aspect of the present application provides a mirror data processing device, including:
[0037] A keyword receiving module is used to receive a keyword sent by any DPU node in a target network composed of multiple DPU nodes that communicate with each other and share data, wherein the keyword includes: an image query keyword and / or a block file query keyword; each block file corresponding to the image data is pre-distributed to each of the DPU nodes for distributed storage;
[0038] A block file information search module, used to use the received keyword as the current target keyword, and search for the block file information corresponding to the target keyword from the correspondence between each keyword and each block file information pre-stored in the target network;
[0039] The block file information sending module is used to return the storage location identifier pointed to by the block file information corresponding to the target keyword to the DPU node that issued the target keyword, so that the DPU node downloads each block file corresponding to the target keyword from other DPU nodes corresponding to the storage location identifier and composes the mirror data.
[0040] The fourth aspect of the present application provides a DPU for executing the mirror data processing method mentioned in the first aspect, or for executing the mirror data processing method mentioned in the second aspect.
[0041] A fifth aspect of the present application provides a mirror processing system, comprising: a target network consisting of a plurality of DPU nodes that communicate with each other and share data;
[0042] Any DPU node in the target network is used to be designated as a mirror service node, or a server that is respectively connected to each of the DPU nodes is used to be designated as the mirror service node;
[0043] The mirror service node is used to execute the mirror data processing method mentioned in the first aspect;
[0044] The DPU node that is not designated as the mirror service node is used for the mirror data processing method mentioned in the second aspect.
[0045] A sixth aspect of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the mirror data processing method mentioned in the first aspect when executing the computer program.
[0046] The seventh aspect of the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the mirror data processing method mentioned in the first aspect.
[0047] The eighth aspect of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the mirror data processing method mentioned in the first aspect, or implements the mirror data processing method mentioned in the second aspect.
[0048] The mirror data processing method provided by the present application receives a keyword sent by any DPU node in a target network composed of multiple DPU nodes that communicate with each other and share data, wherein the keyword includes: a mirror query keyword and / or a block file query keyword; each block file corresponding to the mirror data is pre-distributed to each of the DPU nodes for distributed storage; the received keyword is used as the current target keyword, and the block file information corresponding to the target keyword is searched from the correspondence between each keyword and each block file information pre-stored in the target network; the storage location identifier pointed to by the block file information corresponding to the target keyword is returned to the DPU node that issued the target keyword, so that the DPU node downloads each block file corresponding to the target keyword from other DPU nodes corresponding to the storage location identifier, and forms the mirror data; the mirror warehouse technology can be effectively unloaded to the DPU, the load burden of the host can be effectively reduced, the service quality of the application can be effectively improved, the host can achieve good load isolation, and the network transmission performance of the mirror data can be effectively improved, thereby improving the download efficiency and reliability of the mirror data.
[0049] Additional advantages, purposes, and features of the present application will be partially described in the following description, and will become partially apparent to those skilled in the art after studying the following, or may be learned from the practice of the present application. The purposes and other advantages of the present application can be achieved and obtained by the structures specifically pointed out in the specification and the drawings.
[0050] Those skilled in the art will understand that the purposes and advantages that can be achieved by the present application are not limited to the above specific description, and the above and other purposes that can be achieved by the present application will be more clearly understood based on the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] The drawings described herein are used to provide a further understanding of the present application, constitute a part of the present application, and do not constitute a limitation of the present application. The components in the drawings are not drawn to scale, but are only for the purpose of illustrating the principles of the present application. In order to facilitate the illustration and description of some parts of the present application, the corresponding parts in the drawings may be enlarged, that is, they may become larger relative to other components in the exemplary device actually manufactured according to the present application. In the drawings:
[0052] Figure 1 1 is a schematic diagram of a first flow chart of a first mirror data processing method in an embodiment of the present application.
[0053] Figure 2 This is a second flow chart of the first mirror data processing method in one embodiment of the present application.
[0054] Figure 3 1 is a schematic diagram of a first flow chart of a second mirror data processing method in an embodiment of the present application.
[0055] Figure 4 1 is a second flow chart of a second mirror data processing method in an embodiment of the present application.
[0056] Figure 5 Schematic diagram of the functions of a mirroring service device in one embodiment of the present application.
[0057] Figure 6 This is a functional schematic diagram of a DPU node that executes the second mirror data processing method in one embodiment of the present application.
[0058] Figure 7 It is a schematic diagram of the execution logic of the mirror data processing method in an application example of the present application. DETAILED DESCRIPTION
[0059] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the implementation modes and the accompanying drawings. Here, the illustrative implementation modes and descriptions of the present application are used to explain the present application, but are not intended to limit the present application.
[0060] It should also be noted here that in order to avoid obscuring the present application due to unnecessary details, only the structures and / or processing steps closely related to the scheme according to the present application are shown in the accompanying drawings, while other details that are not very relevant to the present application are omitted.
[0061] It should be emphasized that the term “include / comprises” when used herein refers to the presence of features, elements, steps or components, but does not exclude the presence or addition of one or more other features, elements, steps or components.
[0062] It should also be noted that, unless otherwise specified, the term “connection” herein may refer not only to a direct connection but also to an indirect connection involving an intermediate.
[0063] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. In the accompanying drawings, the same reference numerals represent the same or similar components, or the same or similar steps.
[0064] It should be noted that DPU (Data Processing Unit) refers to data processing unit.
[0065] Currently, mainstream image repositories generally adopt a centralized architecture, where a central server is responsible for managing and distributing images. For example:
[0066] Docker Hub is the most popular Docker image registry, provided by Docker Inc. It is a centralized service for storing and sharing Docker images. Most Docker users use Docker Hub to obtain and share images.
[0067] Quay.io: Quay.io is another centralized Docker image registry provided by RedHat. It provides some advanced features such as automatic building and storage of images.
[0068] Amazon Elastic Container Registry (ECR) is a centralized Docker image registry service provided by Amazon AWS, designed for container image management in the AWS environment.
[0069] Google Container Registry (GCR) is a centralized Docker image registry service provided by Google Cloud Platform for storing and distributing Docker images.
[0070] However, the above existing centralized image repositories have some shortcomings, some of the main problems include:
[0071] Single point of failure: The server in the centralized warehouse architecture is the core of the entire system. If this central server fails or is unavailable, users will not be able to access or share images. This single point of failure may have a significant impact on the production environment.
[0072] Bandwidth bottleneck: When users download or upload images, they need to rely on the bandwidth of the central warehouse. In large-scale or high-traffic situations, the bandwidth of the central server may become a bottleneck, resulting in a decrease in download and upload speeds.
[0073] Network latency: Since users need to communicate with the central server over the Internet, they may be affected by network latency. This may cause problems in distributed teams, deployments across geographical locations, or in environments with unstable networks.
[0074] Security issues: Centralized warehouses may become targets of attacks because they store a large amount of image data. If security is insufficient, they may face potential vulnerabilities and attack risks.
[0075] Dependence on commercial services: Some centralized warehouse services may be provided by specific providers and may require a paid subscription. For some organizations, reliance on commercial services may bring cost and controllability issues.
[0076] Lack of offline support: Centralized warehouses usually need to be connected to the Internet, which may pose challenges for some scenarios in isolated or no network environment. In some cases, such as internal networks in production environments, the lack of offline support may not be ideal.
[0077] In order to solve the above problems, the embodiments of the present application provide a mirror data processing method, a mirror data processing device or DPU for executing the mirror data processing method, a mirror data processing system, etc., and the DPU is used as a hardware unit to accelerate data processing, which can achieve the following effects:
[0078] (1) Network bandwidth and throughput: The DPU has high-performance network features, which can significantly improve network bandwidth and throughput. By offloading the image repository to the DPU, these high-performance network interfaces can be fully utilized to increase the download and upload speed of images.
[0079] (2) Computing acceleration: DPUs are usually equipped with specialized hardware accelerators that can accelerate specific types of computing tasks, such as deep learning reasoning. By working together with the image repository and the DPU, some computing-intensive operations can be performed on the DPU, which can reduce the burden on the host CPU, reduce the impact on the main business, and improve the overall system performance.
[0080] (3) Intelligent caching and distribution: The DPU integrates intelligent caching and distribution image technology to optimize the image caching strategy based on workload and demand, thereby managing and distributing images more efficiently.
[0081] (4) Reduce host load: Offloading the image repository to the DPU can share the burden of the host CPU and memory, improving the host's availability and response speed. This is particularly useful in high-load environments.
[0082] (5) Higher concurrent processing capabilities: DPUs usually have highly parallel processing capabilities. By making full use of their concurrency, the concurrent request processing capabilities of the image repository can be improved.
[0083] In general, offloading image warehouse technology to the DPU mainly solves problems related to network performance, computing acceleration, intelligent caching and distribution, and host resource burden, thereby improving the performance and efficiency of the entire system. This may bring significant benefits especially for large-scale, high-concurrency, and computing-intensive scenarios.
[0084] The details are described in detail through the following examples.
[0085] Based on this, the embodiment of the present application provides a first mirror data processing method that can be implemented by a mirror service node (also referred to as a mirror data processing device, a DPU mirror service end, or a mirror distribution engine, etc.) or any DPU node designated as the mirror service node, see Figure 1 The first mirror data processing method specifically includes the following contents:
[0086] Step 100: In a target network composed of multiple DPU nodes that communicate with each other and share data, receive a keyword sent by any of the DPU nodes, wherein the keyword includes: a mirror query keyword and / or a block file query keyword; each block file corresponding to the mirror data is pre-distributed to each of the DPU nodes for distributed storage.
[0087] Before transmitting between different DPU distribution nodes through a peer-to-peer network, it is necessary to first build a target network with each DPU as a DPU node so that each DPU node can communicate with each other and share data. In one example, the target network can adopt a peer-to-peer network (i.e., a P2P network), which can be built based on a P2P protocol.
[0088] In one or more embodiments of the present application, multiple DPU nodes that communicate with each other and share data may mean that each DPU node communicates and shares data with all other DPU nodes respectively, or that each DPU node communicates and shares data only with the next adjacent node in one-way propagation, or that each DPU node can communicate and share data with at least one other node. The specific settings can be made according to actual needs.
[0089] It is understandable that each block file corresponding to the mirror data can be pre-distributed by the mirror service node to each of the DPU nodes for distributed storage.
[0090] Step 200: The received keyword is used as the current target keyword, and the block file information corresponding to the target keyword is searched from the correspondence between each keyword and each block file information pre-stored in the target network.
[0091] In step 200, the correspondence between each keyword and each block file information pre-stored in the target network can be stored in a data table, for example, a distributed table can be used. If the keyword or subsequent unique identifier refers to a hash value, the data table can directly use a distributed hash table (DHT).
[0092] Specifically, a distributed hash table (DHT) algorithm is used to determine data storage and search in a distributed system, and the data is dispersedly stored on each DPU node in the entire DPU distributed mirror network. DPUs are usually equipped with dedicated hardware accelerators. By offloading DHT calculations, the burden on the host CPU can be reduced, the impact on the main business can be reduced, and the overall system performance can be improved.
[0093] Step 300: Return the storage location identifier pointed to by the block file information corresponding to the target keyword to the DPU node that issued the target keyword, so that the DPU node downloads each block file corresponding to the target keyword from other DPU nodes corresponding to the storage location identifier and forms the mirror data.
[0094] In step 300, the DPU node downloading the block file directly communicates with the node storing the block file DPU to request the required block file. This process does not require the intervention of the central server, but finds the node storing the block by searching the corresponding relationship between each keyword in the target network and each block file information. At the same time, due to the use of the high-performance network characteristics of the DPU, the transmission speed of the block file can be greatly improved.
[0095] From the above description, it can be seen that the first image data processing method provided in the embodiment of the present application can effectively offload the image warehouse technology to the DPU, can effectively reduce the load burden of the host, can effectively improve the service quality of the application, so that the host can achieve load isolation well, and can effectively improve the network transmission performance of the image data, thereby improving the download efficiency and reliability of the image data.
[0096] In order to further improve the efficiency of image data uploading and the reliability and effectiveness of subsequent downloading, in a first image data processing method provided in an embodiment of the present application, see Figure 2 , the first mirror data processing method further specifically includes the following contents before step 100:
[0097] Step 010: Receive mirror data.
[0098] In step 010, the mirror service node receives the mirror data uploaded by the user.
[0099] Step 020: Split the mirror data to obtain multiple block files, and set unique identifiers for the mirror data and each of the block files.
[0100] It is understandable that the unique identifier may be a hash value or a unique code generated based on other encoding methods.
[0101] Step 030: Generate seed information corresponding to each of the block files, wherein the seed information includes a unique identifier of the block file, and also includes the name of the mirror data to which the block file belongs, the size of the mirror data and the unique identifier of the mirror data.
[0102] Step 040: Distribute each of the block files and the corresponding seed information to each of the DPU nodes for distributed storage.
[0103] For each DPU node, the received block files and corresponding seed information are stored accordingly, so that when a certain information in the seed information is received, the corresponding block file can be quickly located. This method can also enable the same DPU node to be used to store multiple block files.
[0104] And, step 050: using the unique identifier of each block file as a block file query keyword, and / or using the unique identifier or name of the mirror data as the mirror query keyword, storing a one-to-one correspondence between the unique identifier corresponding to each of the block files, the name of the mirror data to which it belongs, the size of the mirror data and the unique identifier of the mirror data.
[0105] The one-to-one correspondence between the unique identifier corresponding to each of the block files, the name of the image data to which it belongs, the size of the image data and the unique identifier of the image data can be stored in a distributed hash table. In a distributed hash table, data can be divided into key-value pairs, where the key is a unique identifier generated according to a certain hash algorithm, and the value is the data block or metadata information to be stored. Each image or image block can be regarded as a data item, where the key can be the hash value or other unique identifier of the image, and the value is the content of the image block data.
[0106] Specifically, the image service node can split the uploaded image data into multiple block files (can be referred to as: blocks) and generate seed information. The seed information includes the metadata of the image (image name, size, hash value, hash value of each block), and adds a unique identifier (hash value) to each block file, and distributes and stores them on multiple DPU nodes of the target network. The image service node splits the large image from the image uploaded by the user into multiple blocks (these blocks are different layers (layers) or other logical blocks of the image, where the logical block means: one layer and one block, or one layer is divided into two blocks). Each layer of the image can be used as a separate block file, and then these blocks are transmitted between different DPU distribution nodes through a peer-to-peer network. Each block has a unique identifier (usually a hash value) and can be transmitted independently.
[0107] The transmission can be done in a concurrent manner, and each node will select the seeds to be stored according to certain strategies or rules: based on factors such as network topology, the relationship between nodes, and the popularity of seeds. The distribution and search of seeds are managed by using a data structure similar to DHT (distributed hash table).
[0108] In order to further improve the reliability and efficiency of image data distribution and downloading, in a first image data processing method provided in an embodiment of the present application, see Figure 2 , the first mirror data processing method further includes the following contents before step 010:
[0109] Step 001: Obtain the current network topology type of the target network.
[0110] Step 002: Determine the current propagation path corresponding to the target network according to the network topology type.
[0111] Step 003: Distribute the propagation path to each of the DPU nodes in the target network, so that the DPU node that issues the target keyword searches for other DPU nodes corresponding to the storage location identifier in the target network based on the propagation path when or after receiving the storage location identifier, and downloads each block file corresponding to the target keyword from the other DPU nodes corresponding to the storage location identifier.
[0112] In order to further improve the reliability and efficiency of image data distribution and downloading, in a first image data processing method provided in an embodiment of the present application, see Figure 2 , after step 003 or before or after other steps in the first mirror data processing method, the following contents are specifically included:
[0113] Step 004: If the propagation path shows that there is a DPU node whose adjacent DPU node is not unique, then the DPU node is used as the current target node.
[0114] Step 005: Obtain the current status information of multiple adjacent DPU nodes corresponding to the target node, and based on the current status information of each of the adjacent DPU nodes.
[0115] Step 006: Select one of the adjacent DPU nodes corresponding to the target node as the current neighbor node of the target node.
[0116] Step 007: Send the selection information corresponding to the neighbor node to the target node, so that the target node selects the neighbor node from multiple adjacent DPU nodes.
[0117] In order to further improve the propagation efficiency and the downloading efficiency and reliability of block files, in a first mirror data processing method provided in an embodiment of the present application, the state information includes: at least one of the current bandwidth, load information, location information, availability information and popularity information of the adjacent DPU node;
[0118] The availability information includes: the number and type of block files currently stored by the adjacent DPU node, and the download volume for the stored block files;
[0119] The popularity information includes: the download frequency or the sharing number of the block file currently stored in the adjacent DPU node in the target network.
[0120] Specifically, the Image Service Node can perform the following:
[0121] (1) Obtaining DPU node bandwidth: Each DPU node has upload and download bandwidth limits, which will affect the data transmission speed and priority between nodes. Therefore, the input data of the DPU intelligent distribution engine also needs to include the bandwidth of each node for selecting neighbor nodes.
[0122] (2) Collecting DPU node load information: Obtain monitoring data of the DPU node, including the current load of the DPU node, including the usage of resources such as CPU, memory, disk, and the current network load, for selecting neighbor nodes.
[0123] (3) Input geographic location: Obtain the physical location of the DPU node for selecting neighbor nodes, which affects the network latency and reliability between nodes.
[0124] (4) Maintaining the availability information of data blocks: that is, the number and type of data blocks stored on each DPU node, as well as the requests for these data blocks, for the purpose of selecting neighbor nodes.
[0125] (5) Setting the importance or popularity information of data blocks: Some data blocks may be more important than other data blocks, such as commonly used system files or popular content.
[0126] In order to further improve the reliability and integrity of image distribution, in a first image data processing method provided in an embodiment of the present application, the DPU node that issues the target keyword verifies the integrity and correctness of each block file based on preset verification rules when or after downloading each block file corresponding to the target keyword, and combines each block file after the verification passes to obtain the image data.
[0127] In order to further improve the applicability of mirror data processing, in a first mirror data processing method provided in an embodiment of the present application, the network topology structure type includes: a tree network, a mesh network or a point-to-point network;
[0128] Correspondingly, if the current network topology type of the target network is the tree network, the propagation path corresponding to the tree network includes: propagating from top to bottom in the tree structure formed by each DPU node in the tree network;
[0129] If the current network topology type of the target network is the mesh network, the propagation path corresponding to the mesh network includes: propagation through each adjacent DPU node in the mesh network;
[0130] If the current network topology type of the target network is the point-to-point network, the propagation path corresponding to the mesh network includes: mutual propagation between the DPU nodes that communicate with each other in the point-to-point network.
[0131] Specifically, the mirror service node can first identify the connection relationship between nodes in the target network, such as peer-to-peer networks, tree networks, mesh networks, etc. Different topological structures affect the path and efficiency of information dissemination. Therefore, the mirror service node will obtain the current network topological structure, and then it can perform subsequent processing. For example: for a tree network, data can be propagated from top to bottom through the tree structure; for a mesh network, data can be propagated directly through adjacent nodes; for a peer-to-peer network, data can be propagated between multiple nodes.
[0132] The embodiment of the present application provides a second mirror data processing method that can be implemented by the DPU node that is not designated as the mirror service node, see Figure 3 The second mirror data processing method specifically includes the following contents:
[0133] Step 400: In a target network composed of multiple DPU nodes that communicate with each other and share data, a download request for mirror data is received, and corresponding keywords are obtained from the download request, the keywords including: mirror query keywords and / or block file query keywords; each block file corresponding to the mirror data is pre-distributed to each of the DPU nodes for distributed storage.
[0134] Step 500: Send the keyword to the mirror service node corresponding to the target network, so that the mirror service node uses the received keyword as the current target keyword, and searches for the block file information corresponding to the target keyword from the correspondence between each keyword and each block file information pre-stored in the target network, and then returns the storage location identifier pointed to by the block file information corresponding to the target keyword.
[0135] Step 600: Receive the storage location identifier, and download each block file corresponding to the target keyword from other DPU nodes corresponding to the storage location identifier to form the mirror data.
[0136] From the above description, it can be seen that the second image data processing method provided in the embodiment of the present application can effectively offload the image warehouse technology to the DPU, can effectively reduce the load burden of the host, can effectively improve the service quality of the application, so that the host can achieve load isolation well, and can effectively improve the network transmission performance of the image data, thereby improving the download efficiency and reliability of the image data.
[0137] In order to further improve the efficiency of image data uploading and the reliability and effectiveness of subsequent downloading, in a second image data processing method provided in an embodiment of the present application, see Figure 4 The second mirror data processing method further includes the following contents before step 400:
[0138] Step 050: Receive and store the block file and corresponding seed information sent by the mirror service node, wherein the block file is obtained by the mirror service node after splitting the mirror data it receives, and the mirror service node also sets unique identifiers for the mirror data and each of the block files, and generates seed information corresponding to each of the block files, wherein the seed information includes the unique identifier of the block file, and also includes the name of the mirror data to which the block file belongs, the size of the mirror data and the unique identifier of the mirror data; the mirror service node then distributes each of the block files and the corresponding seed information to each of the DPU nodes for distributed storage.
[0139] Furthermore, the mirror service node uses the unique identifier of each block file as a block file query keyword, and / or uses the unique identifier or name of the mirror data as a mirror query keyword, and stores a one-to-one correspondence between the unique identifier corresponding to each block file, the name of the mirror data to which it belongs, the size of the mirror data, and the unique identifier of the mirror data.
[0140] In order to further improve the reliability and efficiency of image data distribution and downloading, in a second image data processing method provided in an embodiment of the present application, see Figure 4 The second mirror data processing method further includes the following contents before step 050:
[0141] Step 008: Receive the current propagation path corresponding to the target network sent by the mirror service node, and then when or after receiving the storage location identifier, search for other DPU nodes corresponding to the storage location identifier in the target network based on the propagation path, and download each block file corresponding to the target keyword from the other DPU nodes corresponding to the storage location identifier.
[0142] In order to further improve the reliability and integrity of image distribution, in a second image data processing method provided in an embodiment of the present application, see Figure 4 , step 600 in the second mirror data processing method further specifically includes the following contents:
[0143] Step 610: Receive the storage location identifier, and download each block file corresponding to the target keyword from other DPU nodes corresponding to the storage location identifier.
[0144] Step 620: When or after each block file corresponding to the target keyword is downloaded, the integrity and correctness of each block file is verified based on a preset verification rule, and after the verification passes, each block file is combined to obtain the mirror data.
[0145] The present application also provides a mirror data processing device (i.e., the aforementioned mirror service node) for executing all or part of the contents of the first mirror data processing method, see Figure 5 , the mirror data processing device specifically includes the following contents:
[0146] The keyword receiving module 10 is used to receive a keyword sent by any DPU node in a target network composed of multiple DPU nodes that communicate with each other and share data, wherein the keyword includes: a mirror query keyword and / or a block file query keyword; each block file corresponding to the mirror data is pre-distributed to each of the DPU nodes for distributed storage.
[0147] The block file information search module 20 is used to use the received keyword as the current target keyword and search for the block file information corresponding to the target keyword from the correspondence between each keyword and each block file information pre-stored in the target network.
[0148] The block file information sending module 30 is used to return the storage location identifier pointed to by the block file information corresponding to the target keyword to the DPU node that issued the target keyword, so that the DPU node downloads each block file corresponding to the target keyword from other DPU nodes corresponding to the storage location identifier and forms the mirror data.
[0149] The embodiment of the mirror data processing device provided in the present application can be specifically used to execute the processing flow of the embodiment of the first mirror data processing method in the above embodiment. Its functions are not repeated here, and reference can be made to the detailed description of the above first mirror data processing method embodiment.
[0150] The part of the mirror data processing device that performs the first mirror data processing can be completed in the server or client device, and can be executed by the DPU node designated as the mirror service node. The specific selection can be based on the processing capability of the client device or node, as well as the limitations of the user's usage scenario. This application does not limit this. If all operations are completed in the client device, the client device may also include a processor for specific processing of the first mirror data processing.
[0151] The client device may have a communication module (i.e., a communication unit) that can communicate with a remote server to achieve data transmission with the server. The server may include a server on the task scheduling center side, and other implementation scenarios may also include a server on an intermediate platform, such as a server on a third-party server platform that has a communication link with the task scheduling center server. The server may include a single computer device, or a server cluster consisting of multiple servers, or a server structure of a distributed device.
[0152] The server and the client device may communicate with each other using any suitable network protocol, including network protocols that have not yet been developed on the date of filing this application. The network protocols may include, for example, TCP / IP, UDP / IP, HTTP, HTTPS, etc. Of course, the network protocols may also include, for example, RPC (Remote Procedure Call Protocol) and REST (Representational State Transfer) protocols used on top of the above protocols.
[0153] From the above description, it can be seen that the image data processing device provided in the embodiment of the present application can effectively offload the image warehouse technology to the DPU, can effectively reduce the load burden of the host, can effectively improve the service quality of the application, so that the host can achieve load isolation well, and can effectively improve the network transmission performance of the image data, thereby improving the download efficiency and reliability of the image data.
[0154] The present application also provides a DPU node (i.e., a DPU device) for executing all or part of the content of the first image data processing method and / or the second image data processing method, which can effectively unload the image warehouse technology to the DPU, effectively reduce the load burden of the host, and effectively improve the service quality of the application program, so that the host can achieve good load isolation, and can effectively improve the network transmission performance of the image data, thereby improving the download efficiency and reliability of the image data.
[0155] join Figure 6 , the DPU node used to execute the second mirror data processing method may specifically include the following contents:
[0156] The request receiving and keyword extraction module 40 is used to receive a download request for mirror data in a target network composed of multiple DPU nodes that communicate with each other and share data, and obtain corresponding keywords from the download request, wherein the keywords include: mirror query keywords and / or block file query keywords; each block file corresponding to the mirror data is pre-distributed to each of the DPU nodes for distributed storage.
[0157] The keyword sending module 50 is used to send the keyword to the mirror service node corresponding to the target network, so that the mirror service node uses the received keyword as the current target keyword, and searches for the block file information corresponding to the target keyword from the correspondence between each keyword and each block file information pre-stored in the target network, and then returns the storage location identifier pointed to by the block file information corresponding to the target keyword.
[0158] The block file downloading and combining module 60 is used to receive the storage location identifier, and download each block file corresponding to the target keyword from other DPU nodes corresponding to the storage location identifier, and compose the mirror data.
[0159] The embodiment of the DPU node for executing the second mirror data processing method provided in the present application can be specifically used to execute the processing flow of the embodiment of the second mirror data processing method in the above-mentioned embodiment. Its functions are not repeated here, and reference can be made to the detailed description of the above-mentioned embodiment of the second mirror data processing method.
[0160] The present application also provides an embodiment of a mirror processing system, the mirror processing system specifically comprising: a target network composed of a plurality of DPU nodes that communicate with each other and share data;
[0161] Any DPU node in the target network is used to be designated as a mirror service node, or a server that is respectively connected to each of the DPU nodes is used to be designated as the mirror service node;
[0162] The mirror service node is used to execute the embodiment of the aforementioned first mirror data processing method;
[0163] The DPU node that is not designated as the mirror service node is used to execute the aforementioned embodiment of the second mirror data processing method.
[0164] In order to further illustrate the above-mentioned first mirror data processing method and second mirror processing method, taking the target network as a P2P network as an example, the present application also provides a specific application example of a mirror data processing method, which can be executed by a mirror service node or a DPU P2P mirror service end. The P2P mirror warehouse technology hardware is unloaded to the DPU. A dedicated network processor can be used to handle network-related tasks, thereby reducing the burden on the host and improving network transmission performance. If the P2P mirror warehouse technology is placed on the host, the performance may be limited by the processing power of the host and the performance of the network interface. Offloading to the hardware DPU can reduce the load on the server, ensure the high service quality of the application, and avoid the inability to achieve good load isolation on the host.
[0165] As a hardware unit for accelerating data processing, DPU combined with P2P technology can achieve the following effects:
[0166] (1) Network bandwidth and throughput: The DPU has high-performance network features, which can significantly improve network bandwidth and throughput. By offloading the P2P image repository to the DPU, these high-performance network interfaces can be fully utilized to improve the download and upload speed of images.
[0167] (2) Computing acceleration: DPUs are usually equipped with specialized hardware accelerators that can accelerate specific types of computing tasks, such as deep learning reasoning. By working together with the P2P image repository and the DPU, some computing-intensive operations can be performed on the DPU, reducing the burden on the host CPU, reducing the impact on the main business, and improving the overall system performance.
[0168] (3) Intelligent caching and distribution: The DPU integrates intelligent caching and distribution image technology to optimize the image caching strategy based on workload and demand, thereby managing and distributing images more efficiently.
[0169] (4) Reducing host load: Offloading the P2P image repository to the DPU can share the burden of the host CPU and memory, improving the host's availability and response speed. This is particularly useful in high-load environments.
[0170] (5) Higher concurrent processing capabilities: DPUs usually have highly parallel processing capabilities. By making full use of their concurrency, the concurrent request processing capabilities of the P2P image repository can be improved.
[0171] In general, offloading P2P image repository technology to the DPU mainly solves problems related to network performance, computing acceleration, intelligent caching and distribution, and host resource burden, thereby improving the performance and efficiency of the entire system. This may bring significant benefits especially for large-scale, high-concurrency, and computing-intensive scenarios.
[0172] In summary, offloading the P2P image warehouse technology to the DPU has certain advantages, but it also has certain difficulties in technical implementation. The P2P image warehouse needs to process a large amount of data and network communication, which will occupy a large amount of DPU resources. Whether it will cause the performance of other services on the DPU to decline, in order to solve the above problems, in the DPU P2P image warehouse technology, when we maintain the DHT table, mirror service nodes such as tracker servers can comprehensively consider the network load of each DPU node, optimize the order of the neighboring node list, and reduce network congestion and delay.
[0173] See also Figure 7 The mirror data processing method provided by the application example of this application specifically includes the following contents:
[0174] 1. Push the image
[0175] Specifically, the user sends the mirror data to the mirror service node, which may include Figure 7 The proxy and tracker in the system can also be independently set as a device, such as a server or a DPU node.
[0176] 2. Upload block files
[0177] Specifically, add an identifier to the block file: the DPU P2P image server splits the uploaded image into multiple blocks and generates multiple seeds. The seed information includes the metadata of the image (image name, size, hash value, hash value of each block).
[0178] 3. Upload tags
[0179] Specifically, a unique identifier (hash value) is added to each block.
[0180] 4. Block files
[0181] Specifically, the image file is first uploaded to HDFS as central storage.
[0182] The image files are then transferred between nodes in the data center through the P2P protocol. When a node needs an image, it can get the data blocks of the image from HDFS and then share these data blocks with other nodes.
[0183] Nodes can exchange data directly through the P2P distribution system without relying on a single central server. This can improve the transmission efficiency of the image and reduce the load pressure on the central server.
[0184] The block files are distributed and stored on multiple DPU nodes in the P2P network.
[0185] That is, the DPU P2P image server splits the large image from the image uploaded by the user into multiple blocks (these blocks are different layers of the image or other logical blocks, where a logical block means: one layer is one block, or one layer is divided into two blocks), each layer of the image can be used as a separate block file, and then these blocks are transmitted between different DPU distribution nodes through the peer-to-peer network. Each block has a unique identifier (usually a hash value) and can be transmitted independently.
[0186] Storage block files: Use the Distributed Hash Table (DHT) algorithm to determine data storage and search in the distributed system, and store the data in each DPU node in the entire DPU distributed mirror network. DPU is usually equipped with a dedicated hardware accelerator, which reduces the burden on the host CPU by offloading DHT calculations, reduces the impact on the main business, and improves the overall system performance.
[0187] In a distributed hash table, data can be divided into key-value pairs, where the key is a unique identifier generated by a hash algorithm, and the value is the data block or metadata information that needs to be stored. Each image or image block can be regarded as a data item, where the key can be the hash value or other unique identifier of the image, and the value is the content of the image block data.
[0188] Find blocks: When a DPU node needs to download a Docker image, it can use the DHT algorithm to find the node that stores the required block. Specifically, after receiving a user request, it can obtain parameters or keywords from the request, and then search in the distributed hash table through the extracted parameters or keywords. The keywords can be the name, label or identifier of the image, etc. The node searches the DHT through the hash value to find the node that stores the corresponding block. In addition, the DPU integrates intelligent distribution technology to optimize the node selection strategy according to the workload. Specifically, the node can be selected based on the monitoring data of the workload.
[0189] Block distribution: The DPU node that downloads the block file directly communicates with the node that stores the block file DPU to request the required block. This process does not require the intervention of the central server, but finds the node that stores the block through the DHT search in the P2P network. At the same time, due to the use of the high-performance network characteristics of the DPU, the transmission speed of the block is greatly improved.
[0190] Chunk verification and assembly: The downloaded chunks are verified by methods such as md5 to ensure their integrity and correctness. Once all the required chunks are downloaded, they can be assembled sequentially based on their ids to restore the large image.
[0191] 5. Seeds
[0192] Specifically, in a P2P network, a seed is a node that has a complete file or data, and it can share this data with other nodes. Seed nodes usually refer to nodes that have complete Docker image files, and they can share the block data of these image files with other nodes. The transmission can be done in a concurrent manner, and each node will select the seeds to be stored according to certain strategies or rules: based on factors such as network topology, the relationship between nodes, and the popularity of seeds. And a data structure similar to DHT (distributed hash table) is used to manage the distribution and search of seeds.
[0193] 6. Tags
[0194] Specifically, the unique identifier is also sent to the DPU node for storage. The index is a structure used to track the metadata information of the Docker image and its various versions. It records the available images and their corresponding tags, version numbers, seed nodes, and other information. When a client requests to download a Docker image, it usually needs to specify the image tag as a parameter of the request. The engine searches for the corresponding image information in the index based on the tag requested by the client, and provides available seed nodes for the client to download.
[0195] 7. Seeds
[0196] Specifically, the Tracker server also records the information of the Seed Nodes, including the complete files or data they have. This information can help other nodes find suitable Seed Nodes to download files.
[0197] 8. Peer
[0198] Specifically, the Tracker server helps clients find available nodes in the network and realize peer-to-peer data transmission by maintaining node information and processing query requests. It plays a role in coordinating and managing the P2P network, thereby realizing efficient distribution and downloading of files.
[0199] 9. Pull the image
[0200] After the DPU node combines the image data, the user can pull the image data from the DPU node. Specifically, the process of pulling the image includes the following:
[0201] (1) Query the Tracker server: The node will first send a query request to the Tracker server, asking for information about other nodes that have the required image file. The Tracker server will query and return a list of available nodes to the requesting node based on the node information it maintains.
[0202] (2) Select the most suitable seed node: Based on the node list returned by the Tracker, the requesting node will select the most suitable seed node as the download source. Usually, the selection of the most suitable seed node may be based on factors such as node status, availability, and network bandwidth.
[0203] (3) Establishing a connection with the seed node: The requesting node will establish a connection with the selected seed node and send a download request to it. The seed node will respond to the request and start transferring the required Docker image file to the requesting node.
[0204] (4) Data transmission and downloading: Data transmission begins between the requesting node and the seed node. The seed node transmits the data blocks of the Docker image file to the requesting node. The requesting node gradually receives and stores these data blocks until the complete image file is downloaded.
[0205] (5) Inspection and verification: Once all the data blocks of the image file are successfully downloaded, the requesting node will inspect and verify the downloaded image file to ensure that the downloaded image file is complete and not damaged.
[0206] (6) Image decompression and loading: Finally, the requesting node will decompress and load the downloaded image file so that it can be used to run the Docker container.
[0207] Overall, the above distributed mirroring technology makes downloading more decentralized and efficient, and reduces the burden on the central server. Each DPU node can act as a data provider and receiver, thus forming a more decentralized distribution network. This has potential advantages for large-scale deployment, high-concurrency environments, and increased download speeds.
[0208] Based on the above content, the application example of this application also provides a DPU P2P intelligent distribution engine (for block storage, tracker peer list; the main body is: DPU P2P mirror server):
[0209] 1. Identify the network topology: the connection relationship between nodes in the P2P network, such as peer-to-peer network, tree network, mesh network, etc. Different topologies affect the path and efficiency of information dissemination. Therefore, the DPU intelligent distribution engine will obtain the current network topology and then perform subsequent processing. For example, for a tree network, data can be propagated from top to bottom through the tree structure; for a mesh network, data can be propagated directly through adjacent nodes; for a peer-to-peer network, data can be propagated between multiple nodes.
[0210] 2. Get DPU node bandwidth: Each DPU node has upload and download bandwidth limits, which will affect the data transmission speed and priority between nodes. Therefore, the input data of the DPU intelligent distribution engine also needs to include the bandwidth of each node for selecting neighbor nodes.
[0211] 3. Collect DPU node load information: Obtain monitoring data of the DPU node, including the current load of the DPU node, including the usage of resources such as CPU, memory, disk, and the current network load, for selecting neighbor nodes.
[0212] 4. Enter geographic location: Get the physical location of the DPU node for selecting neighbor nodes, which affects the network latency and reliability between nodes.
[0213] 5. Maintain the availability information of data blocks: that is, the number and type of data blocks stored on each DPU node, as well as the requests for these data blocks, for selecting neighbor nodes.
[0214] 6. Set the importance of data blocks: Some data blocks may be more important than other data blocks, such as commonly used system files or popular content.
[0215] 7. Distribution strategy: The strategy that determines how data is distributed among nodes, for example:
[0216] Active push: The node regularly sends data blocks to neighbor nodes according to the DHT table to ensure that the neighbor nodes can obtain popular data in a timely manner.
[0217] Request response: When a node receives a data request from another node, it determines the DPU node to send the request to based on the DHT table, and then chooses whether to respond to the request and send the data block to the requesting node based on the popularity (i.e., heat) of the requested data block and the data it owns.
[0218] Dynamic update: Popularity information may change over time, so the popularity information needs to be updated regularly by the DPU P2P mirror service to ensure that the selected neighbor nodes still have high data block popularity and to adjust the neighbor node selection strategy in time, such as selecting neighbor nodes with popularity.
[0219] 8. Algorithm selection: Specific distribution algorithms, such as popularity-based algorithms:
[0220] Data block selection: Calculate the popularity of data blocks: First, we need to count the popularity of data blocks in the P2P network, that is, calculate the request frequency or the number of times each data block is shared in the entire network. Popularity can be estimated or counted based on historical request records, node feedback information, etc.
[0221] Neighbor node selection: Based on the popularity information of data blocks, nodes with highly popular data blocks are selected as neighbor nodes. The specific selection strategy can be to select nodes with data blocks above a certain popularity threshold, or to select the top nodes as neighbors based on the popularity ranking.
[0222] Data distribution strategy: Once the neighbor nodes are selected, different strategies can be used for subsequent data distribution.
[0223] By splitting large images into multiple blocks on the DPU and using the P2P protocol to establish a DPU peer-to-peer network, efficient distributed image downloading can be achieved. Each block has a unique identifier and is stored on different DPU nodes using a distributed hash table (DHT) algorithm. DPU nodes communicate with each other and share data through intelligent distribution technology, using hardware accelerators and high-performance network features to increase the transmission speed of blocks. Download nodes use DHT to find and communicate directly with nodes storing blocks without the intervention of a central server to achieve block distribution. The downloaded blocks are assembled after verification to ensure their integrity and correctness. This decentralized distribution network has potential advantages for large-scale deployment, high-concurrency environments, and increased download speeds.
[0224] In summary, the application example of this application has the following advantages:
[0225] (1) Block file identification: On the DPU, a large image is split into multiple blocks, which are usually different layers or other logical blocks of the image. A unique identifier is generated for each block, usually a hash value. The hash value is mapped through an index to facilitate quick search.
[0226] (2) Establish a DPU peer-to-peer network: Use the P2P protocol to establish a peer-to-peer network between DPUs so that each DPU node can communicate with each other and share data.
[0227] (3) Storage block files: The distributed hash table (DHT) algorithm is used to determine the storage location of data in the entire DPU distributed mirror network. Each DPU node is responsible for storing a portion of the blocks. The DPU is usually equipped with a dedicated hardware accelerator to offload DHT calculations, reduce the burden on the host CPU, and improve overall system performance.
[0228] (4) Finding blocks: When a node needs to download a Docker image, Tracker can help the node find other nodes that have the required data. Tracker maintains a list of which peer nodes have which data and provides connection information to find the node that stores the corresponding block. DPU integrates intelligent distribution technology to optimize node selection strategy based on workload.
[0229] (5) Block distribution: The downloading node directly communicates with the node storing the block to request the required block. This process does not require the intervention of the central server, but instead finds the node storing the block through the DHT search in the P2P network. By utilizing the high-performance network characteristics of the DPU, the transmission speed of the block is greatly improved.
[0230] (6) Chunk Verification and Assembly: The downloaded chunks are verified to ensure their integrity and correctness. Once all required chunks are downloaded, they can be assembled to restore the entire image.
[0231] The embodiment of the present application also provides an electronic device, which may include a processor, a memory, a receiver and a transmitter, wherein the processor is used to execute the first mirror data processing method and / or the second mirror data processing method mentioned in the above embodiment, wherein the processor and the memory may be connected via a bus or other means, such as by bus connection. The receiver may be connected to the processor and the memory via wired or wireless means.
[0232] The processor may be a central processing unit (CPU). The processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or a combination of the above chips.
[0233] The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer executable programs and modules, such as program instructions / modules corresponding to the mirror data processing method in the embodiment of the present application. The processor executes various functional applications and data processing of the processor by running the non-transitory software programs, instructions and modules stored in the memory, that is, implementing the first mirror data processing method and / or the second mirror data processing method in the above method embodiment.
[0234] The memory may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created by the processor, etc. In addition, the memory may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0235] The one or more modules are stored in the memory, and when executed by the processor, perform the first mirror data processing method and / or the second mirror data processing method in the embodiment.
[0236] In some embodiments of the present application, the user equipment may include a processor, a memory, and a transceiver unit, which may include a receiver and a transmitter. The processor, memory, receiver, and transmitter may be connected through a bus system. The memory is used to store computer instructions, and the processor is used to execute the computer instructions stored in the memory to control the transceiver unit to send and receive signals.
[0237] As an implementation method, the functions of the receiver and the transmitter in the present application can be considered to be implemented through a transceiver circuit or a dedicated chip for transceiver, and the processor can be considered to be implemented through a dedicated processing chip, a processing circuit or a general chip.
[0238] As another implementation method, it is possible to use a general-purpose computer to implement the server provided in the embodiment of the present application, that is, to store the program code for implementing the functions of the processor, receiver, and transmitter in a memory, and the general-purpose processor implements the functions of the processor, receiver, and transmitter by executing the code in the memory.
[0239] The embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the first mirror image data processing method and / or the second mirror image data processing method are implemented. The computer-readable storage medium can be a tangible storage medium, such as a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a floppy disk, a hard disk, a removable storage disk, a CD-ROM, or any other form of storage medium known in the technical field.
[0240] It should be understood by those skilled in the art that the exemplary components, systems and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software or a combination of the two. Whether it is specifically performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link via a data signal carried in a carrier.
[0241] It should be clear that the present application is not limited to the specific configuration and processing described above and shown in the figures. For the sake of simplicity, a detailed description of the known method is omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order between the steps after understanding the spirit of the present application.
[0242] In the present application, features described and / or illustrated for one embodiment may be used in the same manner or in a similar manner in one or more other embodiments, and / or combined with features of other embodiments or replace features of other embodiments.
[0243] The above description is only the preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the embodiments of the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A mirror data processing method, characterized in that: The method is performed by any DPU node designated as a mirror service node in the target network, and includes: Receive mirror data; Splitting the mirror data to obtain a plurality of block files, and setting a unique identifier for the mirror data and each of the block files; Generate seed information corresponding to each of the block files, wherein the seed information includes a unique identifier of the block file, and also includes a name of the mirror data to which the block file belongs, a size of the mirror data, and a unique identifier of the mirror data; Distributing each of the block files and the corresponding seed information to each of the DPU nodes for distributed storage; And, using the unique identifier of each block file as a block file query keyword, and / or using the unique identifier or name of the mirror data as the mirror query keyword, storing a one-to-one correspondence between the unique identifier corresponding to each block file, the name of the mirror data to which it belongs, the size of the mirror data and the unique identifier of the mirror data; In a target network composed of a plurality of DPU nodes that communicate with each other and share data, a keyword sent by any of the DPU nodes is received, wherein the keyword includes: an image query keyword and / or a block file query keyword; each block file corresponding to the image data is pre-distributed to each of the DPU nodes for distributed storage; The received keyword is used as the current target keyword, and the block file information corresponding to the target keyword is searched from the correspondence between each keyword and each block file information pre-stored in the target network; The storage location identifier pointed to by the block file information corresponding to the target keyword is returned to the DPU node that issued the target keyword, so that the DPU node downloads each block file corresponding to the target keyword from other DPU nodes corresponding to the storage location identifier and forms the mirror data.
2. The mirror data processing method according to claim 1, characterized in that: Before receiving the mirror data, the method further includes: Get the current network topology type of the target network; Determine the current propagation path corresponding to the target network according to the network topology type; The propagation path is distributed to each of the DPU nodes in the target network, so that the DPU node that issues the target keyword searches for other DPU nodes corresponding to the storage location identifier in the target network based on the propagation path when or after receiving the storage location identifier, and downloads each block file corresponding to the target keyword from the other DPU nodes corresponding to the storage location identifier.
3. The mirror data processing method according to claim 2, characterized in that: Also includes: If the propagation path shows that there is a DPU node whose adjacent DPU node is not unique, then the DPU node is used as the current target node; Acquire current status information of multiple adjacent DPU nodes corresponding to the target node, and based on the current status information of each of the adjacent DPU nodes; Selecting one of the adjacent DPU nodes corresponding to the target node as a current neighbor node of the target node; The selection information corresponding to the neighbor node is sent to the target node, so that the target node selects the neighbor node from the plurality of adjacent DPU nodes.
4. The mirror data processing method according to claim 3, characterized in that: The state information includes: at least one of the current bandwidth, load information, location information, availability information and popularity information of the adjacent DPU node; The availability information includes: the number and type of block files currently stored by the adjacent DPU node, and the download volume for the stored block files; The popularity information includes: the download frequency or the sharing number of the block file currently stored in the adjacent DPU node in the target network.
5. The mirror data processing method according to claim 1, characterized in that: When or after downloading each block file corresponding to the target keyword, the DPU node that issues the target keyword verifies the integrity and correctness of each block file based on a preset verification rule, and combines each block file after the verification passes to obtain the mirror data.
6. The mirror data processing method according to claim 2, characterized in that: The network topology types include: tree network, mesh network or point-to-point network; Correspondingly, if the current network topology type of the target network is the tree network, the propagation path corresponding to the tree network includes: propagating from top to bottom in the tree structure formed by each DPU node in the tree network; If the current network topology type of the target network is the mesh network, the propagation path corresponding to the mesh network includes: propagation through each adjacent DPU node in the mesh network; If the current network topology type of the target network is the point-to-point network, the propagation path corresponding to the mesh network includes: mutual propagation between the DPU nodes that communicate with each other in the point-to-point network.
7. A mirror data processing method, characterized in that: Executed by a DPU node that is not designated as a mirror service node in a target network, the method comprising: Receive and store the block file and the corresponding seed information sent by the mirror service node, wherein the block file is obtained by the mirror service node after splitting the mirror data it receives, and the mirror service node also sets a unique identifier for the mirror data and each of the block files, and generates seed information corresponding to each of the block files, wherein the seed information includes the unique identifier of the block file, and also includes the name of the mirror data to which the block file belongs, the size of the mirror data and the unique identifier of the mirror data; the mirror service node then distributes each of the block files and the corresponding seed information to each of the DPU nodes for distributed storage; Furthermore, the mirror service node uses the unique identifier of each block file as a block file query keyword, and / or uses the unique identifier or name of the mirror data as a mirror query keyword, and stores a one-to-one correspondence between the unique identifier corresponding to each block file, the name of the mirror data to which it belongs, the size of the mirror data, and the unique identifier of the mirror data; In a target network composed of a plurality of DPU nodes that communicate with each other and share data, a download request for image data is received, and a corresponding keyword is obtained from the download request, the keyword including: an image query keyword and / or a block file query keyword; each block file corresponding to the image data is pre-distributed to each of the DPU nodes for distributed storage; Sending the keyword to the mirror service node corresponding to the target network, so that the mirror service node uses the received keyword as the current target keyword, searches for the block file information corresponding to the target keyword from the correspondence between each keyword and each block file information pre-stored in the target network, and then returns the storage location identifier pointed to by the block file information corresponding to the target keyword; The storage location identifier is received, and each block file corresponding to the target keyword is downloaded from other DPU nodes corresponding to the storage location identifier to form the mirror data.
8. A DPU, characterized in that: Used to execute the mirror data processing method according to any one of claims 1 to 6, or used to execute the mirror data processing method according to claim 7.
9. An image processing system, characterized in that: include: A target network consisting of multiple DPU nodes that communicate with each other and share data; Any DPU node in the target network is used to be designated as a mirror service node, or a server that is respectively connected to each of the DPU nodes is used to be designated as the mirror service node; The mirror service node is used to execute the mirror data processing method according to any one of claims 1 to 6; The DPU node that is not designated as the mirror service node is used to execute the mirror data processing method described in claim 7.
Citation Information
Patent Citations
Mirror image file management method, device and system, computer equipment and storage medium
CN112565325A
Container mirror image deployment method and device, terminal and medium
CN116880951A