Data broadcasting method and device, computer equipment, storage medium and program product

By decomposing the data set into multiple subsets and copying it to the local cache system in parallel, and exchanging data between computing nodes, the performance bottleneck of the global file system is solved, and data broadcast efficiency and performance are improved.

CN120492181APending Publication Date: 2025-08-15CHINESE PEOPLES LIBERATION ARMY UNIT 32051
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510453262.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In artificial intelligence applications, the performance bottleneck problem of global file systems in the prior art seriously affects data broadcasting efficiency, resulting in low data access performance.

Method used

The data set is decomposed into multiple data subsets, copied in parallel to the local cache system of multiple initial computing nodes, and the data subset is copied in parallel between these nodes. The SSD device of the local cache system is used for data exchange, reducing dependence on the global file system.

Benefits of technology

It effectively reduces the pressure and load of the global file system, improves data broadcast efficiency, and makes full use of the SSD performance and network bandwidth on the computing node.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492181A_ABST
    Figure CN120492181A_ABST
Patent Text Reader

Abstract

The invention relates to a data broadcasting method and device, computer equipment, a storage medium and a program product. The method comprises the following steps: decomposing a data set stored in a global file system into a plurality of data subsets according to the number of initial computing nodes; copying the plurality of data subsets from the global file system to a local cache system of a plurality of initial computing nodes in parallel; wherein the data subsets of the local cache systems copied to the plurality of initial computing nodes are different; and copying the data subsets in parallel among the local cache systems of the plurality of initial computing nodes until the local cache system of each initial computing node comprises the data set, and completing data broadcasting. By adopting the method, the influence of the performance bottleneck problem of the global file system on the data broadcasting efficiency can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a data broadcasting method, apparatus, computer equipment, storage medium, and program product. Background Art

[0002] With the rapid development of artificial intelligence (AI) applications, their demand for data access is increasing. As computing scale increases, the amount of data access increases linearly. AI applications are typically Python programs that rely on a variety of library files. Each AI application process needs to read these files in their entirety. In large-scale environments, the amount of data read from library files alone can reach terabytes, placing enormous pressure on storage systems. Furthermore, AI applications rely on datasets for training. The size and quality of datasets significantly impact training results. Datasets are typically large, often consisting of small files. Therefore, accessing these datasets requires both large data volumes and high random access performance.

[0003] Related technologies typically address this issue by configuring a global file system and local cache systems. Data is first placed on the global file system and then copied in parallel from the global file system to the local cache systems. This approach places a heavy load on the global file system, and factors such as the global file system's throughput, metadata performance, and network bandwidth can become bottlenecks, restricting overall performance and resulting in low data broadcast efficiency. Therefore, mitigating the impact of global file system performance bottlenecks on data broadcast efficiency has become a pressing technical challenge. Summary of the Invention

[0004] Based on this, it is necessary to provide a data broadcasting method, apparatus, computer equipment, storage medium and program product that can reduce the impact of global file system performance bottleneck on data broadcasting efficiency in response to the above technical problems.

[0005] In a first aspect, the present application provides a data broadcasting method, comprising:

[0006] Decompose the data set stored in the global file system into multiple data subsets according to the number of initial computing nodes;

[0007] Copying multiple data subsets from the global file system to the local cache systems of the multiple initial computing nodes in parallel; wherein the data subsets copied to the local cache systems of the multiple initial computing nodes are different;

[0008] The data subsets are copied in parallel between the local cache systems of the multiple initial computing nodes until the local cache system of each initial computing node includes the data set, thus completing the data broadcast.

[0009] In one embodiment, copying a subset of data in parallel between local cache systems of a plurality of initial computing nodes includes:

[0010] Control each initial computing node and select the target computing node corresponding to each initial computing node one by one through a hash modulo method; the target computing node is an initial computing node other than each initial computing node;

[0011] Control the local cache system of each initial computing node to read the data subset in the local cache system of the target computing node corresponding to each initial computing node, or write the data subset in the local cache system of each initial computing node to the local cache system of the target computing node corresponding to each initial computing node.

[0012] In one embodiment, controlling the local cache system of each initial computing node to read the data subset in the local cache system of the target initial computing node corresponding to each initial computing node includes:

[0013] For each initial computing node, controlling the local cache system of the initial computing node, and obtaining in parallel the data files of the data subset in the local cache system of the target computing node corresponding to the initial computing node;

[0014] For each data file, compare the data file with a data subset in the local cache system of the initial computing node;

[0015] When the data file exists in the data subset of the local cache system of the initial computing node, skip the data file;

[0016] When the data file does not exist in the data subset of the local cache system of the initial computing node, the data file is read into the local cache system of the initial computing node.

[0017] In one embodiment, the method further includes:

[0018] When there is a new computing node, the new computing node and the multiple initial computing nodes are used as first computing nodes; the data set stored in the global file system is decomposed into multiple first data subsets according to the number of the first computing nodes;

[0019] Comparing each first data subset in parallel with a data subset in a local cache system of a corresponding first computing node;

[0020] When the first data subsets exist, skipping the copying of the first data subsets;

[0021] When each first data subset does not exist, copy each first data subset to a local cache system of the corresponding first computing node;

[0022] The first data subset is copied in parallel between the local cache systems of the first computing nodes until the local cache system of each first computing node includes the data set, thus completing the data broadcast.

[0023] In one embodiment, before decomposing the data set stored in the global file system into a plurality of data subsets according to the number of initial computing nodes, the method further includes:

[0024] Receive a data broadcast operation carrying a data set and identify whether the data set exists in the global file system;

[0025] When the data set does not exist in the global file system, determining capacity availability of multiple initial computing nodes;

[0026] When the capacity of multiple initial computing nodes does not meet the availability requirements, data is deleted from the data set to obtain the target data set, and the deleted data is marked as deleted.

[0027] Decomposing a data set stored in a global file system into a plurality of data subsets according to the number of initial computing nodes includes: decomposing a target data set into a plurality of data subsets according to the number of initial computing nodes;

[0028] The method further includes: after the target data set is broadcasted, marking the target data set as being in a ready-to-use state.

[0029] In one embodiment, the method further includes:

[0030] When a dataset exists in the global file system, the dataset is broadcast;

[0031] When the capacity of multiple initial computing nodes meets the availability requirements, the data set is broadcast.

[0032] In a second aspect, the present application further provides a data broadcasting device, comprising:

[0033] A data set decomposition module, used for decomposing a data set stored in a global file system into multiple data subsets according to the number of initial computing nodes;

[0034] A data subset copy module is used to copy multiple data subsets from the global file system to the local cache systems of multiple initial computing nodes in parallel; wherein the data subsets copied to the local cache systems of the multiple initial computing nodes are different;

[0035] The data exchange module is used to copy data subsets in parallel between the local cache systems of multiple initial computing nodes until the local cache system of each initial computing node includes the data set, thus completing the data broadcast.

[0036] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0037] Decompose the data set stored in the global file system into multiple data subsets according to the number of initial computing nodes;

[0038] Copying multiple data subsets from the global file system to the local cache systems of the multiple initial computing nodes in parallel; wherein the data subsets copied to the local cache systems of the multiple initial computing nodes are different;

[0039] The data subsets are copied in parallel between the local cache systems of the multiple initial computing nodes until the local cache system of each initial computing node includes the data set, thus completing the data broadcast.

[0040] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the following steps:

[0041] Decompose the data set stored in the global file system into multiple data subsets according to the number of initial computing nodes;

[0042] Copying multiple data subsets from the global file system to the local cache systems of the multiple initial computing nodes in parallel; wherein the data subsets copied to the local cache systems of the multiple initial computing nodes are different;

[0043] The data subsets are copied in parallel between the local cache systems of the multiple initial computing nodes until the local cache system of each initial computing node includes the data set, thus completing the data broadcast.

[0044] In a fifth aspect, the present application further provides a computer program product, comprising a computer program, which, when executed by a processor, implements the following steps:

[0045] Decompose the data set stored in the global file system into multiple data subsets according to the number of initial computing nodes;

[0046] Copying multiple data subsets from the global file system to the local cache systems of the multiple initial computing nodes in parallel; wherein the data subsets copied to the local cache systems of the multiple initial computing nodes are different;

[0047] The data subsets are copied in parallel between the local cache systems of the multiple initial computing nodes until the local cache system of each initial computing node includes the data set, thus completing the data broadcast.

[0048] The above-mentioned data broadcast method, apparatus, computer equipment, storage medium, and program product decompose the data set to evenly distribute the data set to the local SSD devices of all initial computing nodes, and then exchange data between the initial computing nodes. When only one data subset is copied from the global file system, the other data subsets are copied from the local SSD devices of the initial computing nodes. Moreover, the amount of data copied from the global file system is only the size of a data subset, thereby offloading most of the pressure of data copying from the global file system to the local SSD devices of the initial computing nodes, effectively reducing the pressure and load of the global file system, greatly alleviating the performance bottleneck problem caused by the global file system, and reducing the impact of the global file system performance bottleneck problem on data broadcast efficiency. When copying data subsets between initial computing nodes, the SSD devices of all initial computing nodes will perform read and write operations in parallel, which can fully utilize the SSD performance and network bandwidth on the initial computing nodes and effectively improve data broadcast efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0050] Figure 1 A diagram illustrating an application environment of a data broadcasting method according to an embodiment;

[0051] Figure 2 1 is a flow chart of a data broadcasting method according to an embodiment;

[0052] Figure 3 1 is a flow chart of copying a subset of data in parallel between local cache systems of multiple initial computing nodes in one embodiment;

[0053] Figure 4 A diagram of a data broadcast architecture for controlling a local cache system of an initial computing node to read a data subset in a local cache system of a target computing node corresponding to the initial computing node in one embodiment;

[0054] Figure 5 A schematic diagram of a process for controlling the local cache system of each initial computing node to read a data subset in the local cache system of a target initial computing node corresponding to each initial computing node in one embodiment;

[0055] Figure 6 This is a diagram of a data broadcast architecture when a computing node is added in one embodiment;

[0056] Figure 7 1 is a flowchart of lifecycle-based data management steps in one embodiment;

[0057] Figure 8 is a structural block diagram of a data broadcasting device in one embodiment;

[0058] Figure 9 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0059] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0060] In the field of artificial intelligence, AI applications rely on a wide variety of Python libraries and datasets. These datasets contain numerous small files. During application runtime, each application process fully loads the Python libraries it relies on and randomly accesses the small files within the datasets. As computing scale increases, the amount of data accessed increases linearly, and the performance requirements for random access become increasingly stringent. This is particularly true during large-scale AI application testing or debugging, where Python library loading can take up significant time despite the short application runtime. This can severely impact debugging progress and significantly reduce computing resource utilization.

[0061] In the related art, data broadcasting tools in distributed systems include open source software such as mpifileutils, as well as global file systems and local cache systems. While mpifileutils provides a series of parallel file operation tools designed to optimize file management tasks in high-performance computing environments, it primarily targets single file processing and lacks effective support for batch processing of multiple files. In the application model of a global file system and local cache system, the most direct requirement is to distribute data to the local cache systems of compute nodes. The traditional approach is to first place the data on the global file system and then copy the data from the global file system to the compute nodes in parallel. This approach has serious shortcomings, placing a heavy load on the global file system, and the performance of the global file system can also severely restrict the performance of the copy. Furthermore, with the increase in AI applications and the upgrade of AI application algorithms, applications, Python libraries, and datasets are constantly changing, requiring frequent data updates. This places increasingly stringent performance demands on data broadcasting, and the data in the local cache system is also increasing, ultimately leading to insufficient local cache system capacity.

[0062] To address the above problems, a read-only data broadcasting method for large-scale local cache systems is proposed.

[0063] The data broadcasting method provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown in the figure, the terminal 102 communicates with the distributed system 104 through the network to implement data broadcasting. Data broadcasting refers to the technology of sending data from one node to multiple nodes simultaneously. The core goal is to efficiently distribute data to multiple recipients through a single sending operation. Data broadcasting is widely used in scenarios such as large-scale distributed systems and real-time data synchronization. The distributed system 104 includes a global computing node 1042 and multiple initial computing nodes 1044. The global computing node 1042 is configured with a global file system. The global file system refers to a distributed file system that provides global shared access services for data across computing nodes. The initial computing node 1044 is configured with a local cache system. The local cache system can also be called an SSD (Solid State Drive) device local to the initial computing node 1044. It is a mechanism for temporarily storing data in a computing system. It aims to reduce access latency, improve system performance, and reduce the load on back-end storage by storing frequently accessed data locally or close to computing resources.

[0064] Furthermore, a data broadcast program is pre-running in terminal 102, and data broadcasting is accomplished through the data broadcast program. Specifically, the data broadcast program decomposes the data set stored in the global file system into multiple data subsets based on the number of initial computing nodes, and copies the multiple data subsets from the global file system in parallel to the local cache systems of the multiple initial computing nodes. The data subsets copied to the local cache systems of the multiple initial computing nodes are different. The data subsets are copied in parallel between the local cache systems of the multiple initial computing nodes until the data set is included in the local cache system of each initial computing node, completing the data broadcast. Terminal 102 may be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices may include smart speakers, smart TVs, smart air conditioners, smart car devices, and the like. Portable wearable devices may include smart watches, smart bracelets, head-mounted devices, and the like. Global computing node 1042 and initial computing node 1044 may be implemented using independent servers that can provide CPU or GPU computing power, or using a server cluster consisting of multiple servers.

[0065] In an exemplary embodiment, Figure 2 As shown, a data broadcast method is provided, which is applied to Figure 1 The terminal in FIG is taken as an example to illustrate the method, including the following steps 202 to 206. Among them:

[0066] Step 202 : Decompose the data set stored in the global file system into multiple data subsets according to the number of initial computing nodes.

[0067] The data set refers to the data set that needs to be broadcast.

[0068] Optionally, a data broadcast program is running on the terminal. The data broadcast program decomposes the dataset stored in the global computing nodes according to the number of initial computing nodes to generate multiple data subsets. The files included in the data subsets do not overlap, and the number of data subsets generated is determined by the number of initial computing nodes.

[0069] Furthermore, the data set is decomposed according to its characteristics. Specifically, one way is to divide the data set equally according to the number of files in the data set. First, scan the data set, count the number of all files in the data set, and then divide the number of files equally according to the number of initial computing nodes that need to be broadcast, to obtain multiple data subsets with equal number of files. Another way is to divide the data set equally according to the data size of the data set. First, scan the data set, count the total size of all files in the data set, and then divide the file size equally according to the number of initial computing nodes that need to be broadcast, to obtain multiple data subsets with equal file size. The principle of data set decomposition is to ensure that the performance of the generated data subsets copied from the global file system to the initial computing nodes remains as consistent as possible.

[0070] Step 204 : copying the multiple data subsets from the global file system to the local cache systems of the multiple initial computing nodes in parallel; wherein the data subsets copied to the local cache systems of the multiple initial computing nodes are different.

[0071] Different data subsets are copied from the global file system to different initial compute nodes. The copying process is performed in parallel across all initial compute nodes, and also for each initial compute node. After each initial compute node completes the copy, it waits synchronously until all initial compute nodes have completed copying the data subset. After the copying is complete, each initial compute node has a different copy of the data subset, and all data subsets are combined to form a complete data set.

[0072] Step 206 : Copying the data subsets in parallel between the local cache systems of the multiple initial computing nodes until the local cache system of each initial computing node includes the data set, thus completing the data broadcast.

[0073] The initial computing nodes copy the data subsets in parallel. Specifically, through SSD device collaboration, that is, local cache system collaboration, the data subsets are broadcast from the SSD device to other SSD devices until the local cache system of each initial computing node contains the complete data set, at which point the data broadcast process is complete.

[0074] The above-mentioned parallel data broadcasting method decomposes the data set to evenly distribute the data set to the local SSD devices of all initial computing nodes, and then exchanges data between the initial computing nodes. When only one data subset is available, it will be copied from the global file system, and the other data subsets will be copied from the local SSD devices of the initial computing nodes. Moreover, the amount of data copied from the global file system is only the size of a data subset, thereby offloading most of the pressure of data copying from the global file system to the local SSD devices of the initial computing nodes, effectively reducing the pressure and load of the global file system, greatly alleviating the performance bottleneck problem brought by the global file system, and reducing the impact of the global file system performance bottleneck problem on data broadcast efficiency. When copying data subsets between initial computing nodes, the SSD devices of all initial computing nodes will perform read and write operations in parallel, which can fully utilize the SSD performance and network bandwidth on the initial computing nodes and effectively improve the efficiency of data broadcasting.

[0075] In an exemplary embodiment, Figure 3 As shown, in step 206, performing data subset copying in parallel between the local cache systems of the multiple initial computing nodes includes steps 302 to 304.

[0076] Step 302 : Control each initial computing node and select the target computing node corresponding to each initial computing node one by one through a hash modulo method; the target computing node is an initial computing node other than each initial computing node.

[0077] Step 304: Control the local cache system of each initial computing node to read the data subset in the local cache system of the target computing node corresponding to each initial computing node, or write the data subset in the local cache system of each initial computing node to the local cache system of the target computing node corresponding to each initial computing node.

[0078] The data broadcast program controls the local cache systems of multiple initial computing nodes to copy data subsets in parallel. Specifically, for each initial computing node, the target computing node of the initial computing node is selected one by one by hash modulo method. The local cache system of the initial computing node is controlled to read the data subset in the local cache system of the target computing node corresponding to the initial computing node, or the data subset in the local cache system of the initial computing node is written to the local cache system of the target computing node corresponding to the initial computing node. In the process of controlling the local cache system of the initial computing node to read the data subset in the local cache system of the target computing node corresponding to the initial computing node, or the data subset in the local cache system of the initial computing node is written to the local cache system of the target computing node corresponding to the initial computing node, only one data subset is read or written at a time, and the reading or writing process of all the initial computing nodes is carried out in parallel. For example, when broadcasting a data set to N initial computing nodes, it is necessary to read or write N-1 data subsets. After each read or write is completed, synchronization is performed, and then the next data subset is read or written. After completing N-1 data subset reads or writes, the local SSD devices of the N initial computing nodes will have a complete data set. After all initial computing nodes are synchronized, the data set broadcast is completed.

[0079] like Figure 4 Figure 2 shows a data broadcast architecture diagram for controlling the local cache system of an initial compute node to read a data subset from the local cache system of a target compute node corresponding to the initial compute node. D1, D2, and D3 represent data subsets, respectively. Compute Node 1, Compute Node 2, and Compute Node 3 all represent initial compute nodes.

[0080] Initially, each initial computing node has only one data subset, that is, computing node 1, computing node 2, and computing node 3 include data subset D1, data subset D2, and data subset D3 respectively. When reading the data subsets for the first time, computing node 1 reads data subset D2 from computing node 2, computing node 2 reads data subset D3 from computing node 3, and computing node 3 reads data subset D1 from computing node 1; when reading the data subsets for the second time, computing node 1 reads data subset D3 from computing node 3, computing node 2 reads data subset D1 from computing node 1, and computing node 3 reads data subset D2 from computing node 2. Figure 4In the example above, after two reads, all computing nodes have three different data subsets, which can form a complete data set. In the above process, each initial computing node reads a data subset from other initial computing nodes at a time, and another initial computing node also reads a data subset from it at the same time. The reading tasks of all initial computing nodes are performed in parallel. Therefore, when broadcasting a data set to N initial computing nodes, N-1 data subset reads are required. After each read is completed, a synchronization is performed before the next data subset is read. After completing N-1 data subset reads, the local computer has the complete data set. After all initial computing nodes are synchronized, the data set broadcast is completed.

[0081] Optionally, a single data subset contains multiple files, and each initial computing node also adopts a parallel approach when reading a data subset. The reading method can be an efficient transmission protocol such as rsync (Remote Synchronization) and tftp (Trivial File Transfer Protocol).

[0082] In this embodiment, since hash modulo only requires a single hash calculation and modulo operation, the time complexity is O(1), making it lighter than other load balancing strategies and suitable for large-scale task scheduling. When copying data subsets between the local cache systems of multiple initial computing nodes, the local cache systems of all computing nodes perform read and write operations in parallel, fully utilizing the performance of the local cache systems on the computing nodes and the network bandwidth, greatly improving data broadcast efficiency.

[0083] In an optional manner of the above embodiment, controlling the local cache system of each initial computing node to read the data subset in the local cache system of the target initial computing node corresponding to each initial computing node includes:

[0084] For each initial computing node, the local cache system of the initial computing node is controlled to obtain in parallel the data files of the data subset in the local cache system of the target computing node corresponding to the initial computing node; for each data file, the data file is compared with the data subset in the local cache system of the initial computing node; when the data file exists in the data subset in the local cache system of the initial computing node, the copying of the data file is skipped; when the data file does not exist in the data subset in the local cache system of the initial computing node, the data file is read into the local cache system of the initial computing node.

[0085] Optionally, the local cache system of each initial computing node can adopt a remote parallel reading method to copy the data files of the data subset in the local cache system of the corresponding target initial computing node. Before reading each data file, a comparison will be performed. Specifically, for each initial computing node, the local cache system of the initial computing node is controlled to obtain in parallel all data files of the data subset in the local cache system of the target computing node corresponding to the initial computing node. For each data file, the data file must be compared with the data subset in the local cache system of the initial computing node. If the data file exists in the data subset, the data file is skipped. If the data file does not exist in the data subset, the actual reading process is carried out to read the data file into the local cache system of the initial computing node, thereby realizing the incremental update of the data in the initial computing node.

[0086] In this embodiment, for each initial computing node, the local cache system of the initial computing node is controlled to concurrently retrieve data files for a data subset from the local cache system of the target computing node corresponding to the initial computing node, effectively improving data acquisition efficiency. Before reading a data file, the data file is compared with the data subset in the local cache system of the initial computing node, proactively identifying the existence of files within the data subset. If the data file exists in the data subset of the local cache system of the initial computing node, the data file is skipped, eliminating the need for repeated data reading. This enables incremental data updates, improves data reading efficiency, and effectively handles incremental updates of data sets.

[0087] In an optional manner of the above embodiment, if the computing resources of the initial computing node are limited, the initial computing node can be controlled to sequentially obtain the data files of the data subset in the local cache system of the corresponding target computing node. Figure 5 As shown, in step 304, controlling the local cache system of each initial computing node to read the data subset of the local cache system of the target initial computing node corresponding to each initial computing node includes:

[0088] Step 502: Control the local cache system of each initial computing node to obtain the current data file in the data subset in the local cache system of the target computing node corresponding to each initial computing node.

[0089] Step 504 : Identify whether the data subset in the local cache system of each initial computing node exists and is consistent with the current data file. If so, execute step 506 ; if not, execute step 508 .

[0090] Step 506: Skip the current data file and proceed to step 510.

[0091] Step 508 , read the current data file into the local cache system of each initial computing node. After the reading is completed, execute step 510 .

[0092] Step 510 determines whether the next data file in the data subset of the local cache system of the target initial computing node corresponding to each initial computing node has been obtained. If so, the next data file is updated as the current data file, and the process returns to step 502 to continue reading the data file. If not, step 512 is executed.

[0093] Step 512: The data subset in the local cache system of each initial computing node is read.

[0094] In this embodiment, by controlling the initial computing node to sequentially obtain data files of a data subset from the local cache system of the corresponding target computing node, the data broadcast speed can be improved when the computing resources of the initial computing node are limited.

[0095] In an exemplary embodiment, the method further includes: a data broadcasting step when a computing node is newly added, the step including:

[0096] When there are new computing nodes, the new computing nodes and multiple initial computing nodes are used as first computing nodes; the data set stored in the global file system is decomposed into multiple first data subsets according to the number of first computing nodes; each first data subset is compared in parallel with the data subset in the local cache system of the corresponding first computing node; when each first data subset exists, the copying of each first data subset is skipped; when each first data subset does not exist, each first data subset is copied to the local cache system of the corresponding first computing node; the first data subsets are copied in parallel between the local cache systems of the first computing nodes until the local cache system of each first computing node includes the data set, thereby completing the data broadcast.

[0097] Optionally, after the data subset copy is completed between multiple initial computing nodes, the local cache system of each initial computing node includes a complete data set. When there is a new computing node, since the data subset does not exist in the new computing node, it is necessary to perform parallel data broadcast on the new computing node and multiple initial computing nodes. Specifically, the new computing node and the multiple initial computing nodes are used as first computing nodes, and the data set in the global file system is decomposed into multiple first data subsets according to the number of first computing nodes. The parallel data broadcast process is performed on the multiple first data subsets according to the aforementioned data broadcast method until the local cache system of each first computing node includes a complete data set, indicating that the data broadcast process is completed.

[0098] Each data file in each first data subset is compared in parallel with the data subset in the local cache system of the first computing node corresponding to each first data subset. The first computing node corresponding to each first data subset refers to the storage node of the first data subset. If each data file exists in the data subset in the local cache system of the first computing node corresponding to each first data subset, copying of each data file is skipped. If each data file does not exist in the data subset in the local cache system of the first computing node corresponding to each first data subset, each data file in each first data subset is copied to the local cache system of the corresponding first computing node. The first data subset is then copied in parallel between the local cache systems of the first computing nodes until the data set is included in the local cache system of each first computing node, completing the data broadcast. Because the initial computing nodes contain the complete data set, when a new computing node is added, the initial computing nodes only perform the data comparison process and do not perform the actual copy process. However, the newly added computing nodes must perform the entire actual copy process. Therefore, when performing parallel data broadcasting on M computing nodes with existing datasets and N computing nodes without datasets (newly added computing nodes), the newly added computing nodes only need to read N / (M+N) of the dataset size from the global file system. The remaining datasets are copied from the local SSD devices of other computing nodes. The total read and write data volume is the size of N datasets.

[0099] like Figure 6 The following is a diagram of the data broadcast architecture when a new computing node is added. D1, D2, D3, and D4 represent the first data subset respectively. Computing node 4 is a new computing node, and computing nodes 1, 2, 3, and 4 are used as the first computing node. Initially, computing nodes 1, 2, and 3 already include a complete data set (not shown in the figure). After the new computing node appears, computing nodes 1, 2, and 3 only perform the data comparison process, as shown in the figure. Figure 6 As shown by the dotted line in the middle. Computing node 4 performs the actual copy process, as shown in Figure 6 This is shown by the solid line in the middle. Alternatively, the first data subset D4 corresponding to compute node 4 is copied to compute node 4, and the newly added compute node is controlled to select target nodes one by one through a hash modulo method. The target node is the first compute node other than the newly added compute node. The local cache system of the newly added compute node is controlled to read the data subsets in the local cache systems of each target node until the newly added compute node includes the data set, completing the data broadcast.

[0100] In this embodiment, when there are new computing nodes, the data set in the global file system is decomposed into multiple first data subsets according to the number of new computing nodes and multiple initial computing nodes. Before copying the data subsets, each first data subset is compared in parallel with the data subset in the local cache system of the corresponding first computing node. If each first data subset exists, the copying of each first data subset is skipped, otherwise the actual copying process is performed, thereby actively identifying the existence of the first data subset and being able to effectively deal with the situation of newly added computing nodes.

[0101] In an exemplary embodiment, the method further includes: a data broadcast step when an incremental update occurs to a dataset, which includes: when an incremental update occurs to a dataset in a global file system, decomposing the updated dataset into multiple second data subsets based on the number of computing nodes; comparing the second data subsets with the data subsets in the local cache system of the corresponding computing node; when each second data subset does not exist, copying each second data subset to the local cache system of the corresponding computing node; and copying the second data subsets in parallel between the local cache systems of multiple computing nodes until the updated dataset is included in the local cache system of each computing node, thereby completing the data broadcast. The data broadcast processing process for incremental updates to a dataset is the same as the data broadcast process in the above embodiment and will not be repeated here.

[0102] In an exemplary embodiment, Figure 7 As shown, before decomposing the data set stored in the global file system into multiple data subsets according to the number of initial computing nodes, the method further includes: a lifecycle-based data management step, which includes:

[0103] Step 702: Receive a data broadcast operation carrying a data set.

[0104] Step 704: Identify whether the data set exists in the global file system. If the data set does not exist in the global file system, execute step 706; if the data set exists in the global file system, execute step 708.

[0105] Step 706: Determine the capacity availability of the multiple initial computing nodes. If the capacity of the multiple initial computing nodes does not meet the availability requirement, execute step 710; if the capacity of the multiple initial computing nodes meets the availability requirement, execute step 708.

[0106] Step 708: Broadcast data and then proceed to step 712.

[0107] Step 710 , perform data deletion processing on the data set to obtain the target data set, and mark the deleted data as deleted, and execute step 708 .

[0108] Step 712: After the data broadcast is completed, it is marked as ready for use.

[0109] After the target dataset is broadcasted, the target dataset is marked as ready for use.

[0110] Optionally, a lifecycle-based data management method is an important means of assisting data broadcasting, which is used to record the complete process of data creation, destruction, and use. After receiving a data broadcast operation carrying a data set, it is identified whether the data set exists in the global file system. If it exists, it directly enters the parallel data broadcast process. If it does not exist, it further identifies whether the capacity of multiple initial computing nodes meets the availability requirements. Meeting the availability requirements means that the size of the data set is less than or equal to the capacity of the initial computing node. If the capacity meets the requirements, the parallel data broadcast process is carried out. If the capacity does not meet the requirements, the data elimination mechanism is executed. According to the historical operation mark of the data set, a part of the data set is selected for deletion, and the part of the data is marked as deleted, that is, the pop state, and then enters the parallel data broadcast process. Among them, the data elimination mechanism can be the principle of the least recent use, the least frequent use, etc. After the data broadcast is completed, the data is marked as a state to be used, that is, the push state.

[0111] In this embodiment, by managing the lifecycle of datasets, the generation of large amounts of junk data can be effectively avoided. When storage space is insufficient, sufficient storage space can be provided for new datasets through a data elimination mechanism. This provides a traceable and reliable data management system, providing better storage resource guarantees for AI applications that truly require it. Simultaneously, combined with data broadcasting methods, this can effectively address the problem of insufficient storage space. Furthermore, by recording the usage of datasets in the local cache system, delayed deletion can be used to increase the probability of dataset reuse.

[0112] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0113] Based on the same inventive concept, embodiments of the present application also provide a data broadcasting device for implementing the aforementioned data broadcasting method. The implementation solution provided by this device is similar to the implementation solution described in the aforementioned method. Therefore, the specific limitations in one or more data broadcasting device embodiments provided below can be found in the above-described limitations on the data broadcasting method and will not be further elaborated here.

[0114] In an exemplary embodiment, Figure 8 As shown, a data broadcasting device is provided, comprising: a data set decomposition module 802, a data subset copy module 804 and a data exchange module 806, wherein:

[0115] The data set decomposition module 802 is configured to decompose the data set stored in the global file system into multiple data subsets according to the number of initial computing nodes.

[0116] The data subset copy module 804 is configured to copy multiple data subsets from the global file system to the local cache systems of multiple initial computing nodes in parallel; wherein the data subsets copied to the local cache systems of the multiple initial computing nodes are different.

[0117] The data exchange module 806 is used to copy the data subsets in parallel between the local cache systems of the multiple initial computing nodes until the local cache system of each initial computing node includes the data set, thus completing the data broadcast.

[0118] In an exemplary embodiment, the data exchange module 806 is also used to control each initial computing node, and select the target computing node corresponding to each initial computing node one by one through a hash modulo method; the target computing node is an initial computing node other than each initial computing node; and control the local cache system of each initial computing node to read a data subset in the local cache system of the target computing node corresponding to each initial computing node, or write a data subset in the local cache system of each initial computing node to the local cache system of the target computing node corresponding to each initial computing node.

[0119] In an exemplary embodiment, the data exchange module 806 is also used to control the local cache system of each initial computing node, and in parallel obtain the data files of the data subset in the local cache system of the target computing node corresponding to the initial computing node; for each data file, compare the data file with the data subset in the local cache system of the initial computing node; when the data file exists in the data subset in the local cache system of the initial computing node, skip the data file; when the data file does not exist in the data subset in the local cache system of the initial computing node, read the data file to the local cache system of the initial computing node.

[0120] In an exemplary embodiment,

[0121] The data set decomposition module 802 is also used to, when there is a new computing node, use the new computing node and multiple initial computing nodes as first computing nodes; and decompose the data set stored in the global file system into multiple first data subsets according to the number of first computing nodes.

[0122] The data subset copy module 804 is also used to compare each first data subset in parallel with the data subset in the local cache system of the corresponding first computing node; when each first data subset exists, the copying of each first data subset is skipped; when each first data subset does not exist, the first data subset is copied to the local cache system of the corresponding first computing node.

[0123] The data exchange module 806 is further configured to copy the first data subset in parallel between the local cache systems of the first computing nodes until the local cache system of each first computing node includes the data set, thus completing the data broadcast.

[0124] In an exemplary embodiment, the apparatus further comprises:

[0125] The data cycle management module is used to receive data broadcast operations carrying data sets and identify whether the data sets exist in the global file system. If the data sets do not exist in the global file system, the module determines the capacity availability of multiple initial computing nodes. If the capacity of multiple initial computing nodes does not meet the availability requirements, the module deletes the data sets to obtain the target data sets and marks the deleted data as deleted.

[0126] The data set decomposition module 802 is further configured to decompose the target data set into a plurality of data subsets according to the number of initial computing nodes;

[0127] The data cycle management module is further configured to mark the target data set as being ready for use after the target data set is broadcasted.

[0128] In an exemplary embodiment, the data cycle management module is further configured to broadcast the data set when the data set exists in the global file system; and to broadcast the data set when the capacity of the plurality of initial computing nodes meets the availability requirement.

[0129] Each module in the above-mentioned data broadcasting device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.

[0130] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as shown in FIG. Figure 9 As shown. The computer device includes a processor, memory, an input / output interface, a communication interface, a display unit, and an input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals via wired or wireless means, and the wireless means can be implemented via Wi-Fi, mobile cellular networks, NFC (near field communication), or other technologies. When executed by the processor, the computer program implements a data broadcasting method. The display unit of the computer device is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device casing, or an external keyboard, touchpad or mouse.

[0131] Those skilled in the art will understand that Figure 9 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0132] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.

[0133] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0134] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0135] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0136] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), data processing logic devices based on quantum computing, and the like.

[0137] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0138] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A data broadcasting method, characterized in that: The method comprises: Decompose the data set stored in the global file system into multiple data subsets according to the number of initial computing nodes; Copying multiple data subsets from the global file system to the local cache systems of the multiple initial computing nodes in parallel; wherein the data subsets copied to the local cache systems of the multiple initial computing nodes are different; The data subsets are copied in parallel between the local cache systems of the multiple initial computing nodes until the local cache system of each initial computing node includes the data set, thus completing the data broadcast.

2. The method according to claim 1, characterized in that The copying of data subsets in parallel between the local cache systems of the multiple initial computing nodes includes: Controlling each initial computing node, and selecting a target computing node corresponding to each initial computing node one by one through a hash modulo method; the target computing node is an initial computing node other than each initial computing node; Control the local cache system of each initial computing node to read the data subset in the local cache system of the target computing node corresponding to each initial computing node, or write the data subset in the local cache system of each initial computing node to the local cache system of the target computing node corresponding to each initial computing node.

3. The method according to claim 2, characterized in that The controlling the local cache system of each initial computing node to read the data subset in the local cache system of the target initial computing node corresponding to each initial computing node includes: For each initial computing node, controlling the local cache system of the initial computing node, and obtaining in parallel the data files of the data subset in the local cache system of the target computing node corresponding to the initial computing node; For each data file, compare the data file with a data subset in the local cache system of the initial computing node; When the data file exists in the data subset of the local cache system of the initial computing node, skip the data file; When the data file does not exist in the data subset of the local cache system of the initial computing node, the data file is read into the local cache system of the initial computing node.

4. The method according to any one of claims 1 to 3, characterized in that The method further comprises: When there is a new computing node, the new computing node and the multiple initial computing nodes are used as first computing nodes; the data set stored in the global file system is decomposed into multiple first data subsets according to the number of the first computing nodes; Comparing each first data subset in parallel with a data subset in a local cache system of a corresponding first computing node; When the first data subsets exist, skipping the copying of the first data subsets; When each first data subset does not exist, copy each first data subset to a local cache system of the corresponding first computing node; The first data subset is copied in parallel between the local cache systems of the first computing nodes until the local cache system of each first computing node includes the data set, thus completing the data broadcast.

5. The method according to claim 1, wherein Before decomposing the data set stored in the global file system into a plurality of data subsets according to the number of initial computing nodes, the method further includes: receiving a data broadcast operation carrying a data set, and identifying whether the data set exists in the global file system; When the data set does not exist in the global file system, determining capacity availability of a plurality of initial computing nodes; When the capacity of the multiple initial computing nodes does not meet the availability requirement, data deletion processing is performed on the data set to obtain the target data set, and the deleted data is marked as deleted; Decomposing the data set stored in the global file system into a plurality of data subsets according to the number of initial computing nodes comprises: decomposing the target data set into a plurality of data subsets according to the number of initial computing nodes; The method further includes: after the target data set is broadcasted, marking the target data set as being in a ready-to-use state.

6. The method according to claim 5, characterized in that The method further comprises: When the data set exists in the global file system, broadcasting the data set; When the capacities of the multiple initial computing nodes meet the availability requirement, the data set is broadcasted.

7. A data broadcasting device, characterized in that: The device comprises: A data set decomposition module, used for decomposing a data set stored in a global file system into multiple data subsets according to the number of initial computing nodes; a data subset copy module, configured to copy multiple data subsets from the global file system to the local cache systems of the multiple initial computing nodes in parallel; wherein the data subsets copied to the local cache systems of the multiple initial computing nodes are different; The data exchange module is used to copy data subsets in parallel between the local cache systems of multiple initial computing nodes until the local cache system of each initial computing node includes the data set, thereby completing the data broadcast.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.