Methods, Devices, and Storage Media for Data Management and Distribution in a Machine Learning Platform

By adopting Git version control system and two-level storage structure in the data management and distribution of the machine learning platform, combined with Argo Workflow's task DAG scheduling, the problem of low data transmission efficiency in multiple versions is solved, and data storage and transmission efficiency is improved.

CN119248201BActive Publication Date: 2025-05-30BEIJING INBO DIGITAL TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411569944.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-05
Publication Date
2025-05-30
Estimated Expiration
2044-11-05

AI Technical Summary

Technical Problem

When managing and distributing data from machine learning platforms, the prior art faces the problem of multiple cross-cluster transmission operations when multiple versions of data are used simultaneously, resulting in duplicate data storage work and low transmission efficiency.

Method used

The Git version control system is used to manage and share data. By setting up a two-level storage structure in each Availability Zone, the primary storage is used to mirror the git source station, the secondary storage is used to store file according to the commit id, and the task DAG scheduling is used to realize the sharing of storage space and distribution operations.

Benefits of technology

After a cross-cluster data transmission, it is distributed into multiple versions through secondary storage in the training cluster, avoiding the network load caused by multiple cross-cluster transmission operations. Multiple tasks can share the same version of data, reducing duplicate data storage work, thereby improving data storage and transmission efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119248201B_ABST
    Figure CN119248201B_ABST
Patent Text Reader

Abstract

The present application relates to a method, device, and storage medium for data management and distribution in a machine learning platform. The method includes: managing and sharing data based on a Git version control data system. The data storage includes multiple machine learning training cluster availability zones, and each availability zone is provided with a two-level storage structure. Among them, the primary storage is used to mirror the git source station; the secondary storage is used to store files according to the commit id, and each data version is stored in the secondary storage, and each version has a corresponding commit id. Different training tasks share and call the secondary storage. The data is managed and shared through the Git version control system. When facing the requirements of multiple training tasks, based on the custom resource CR under the Kubernetes platform, directly call the version data required by the task, realize the sharing of storage space and distribution operations among multiple tasks, and effectively improve the data storage and transmission efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of machine learning, and in particular, to a method, device, and storage medium for data management and distribution in a machine learning platform. Background Art

[0002] With the rapid development of AI technologies represented by large models, the volume of training data and model data has increased rapidly. How to effectively manage data, including precise version control, efficient data distribution, and convenient training use, has become a new challenge for AIInfra technology.

[0003] In the existing technologies, when multiple versions of data are in use simultaneously, multiple cross-cluster transmission operations are often required, resulting in duplicate data storage work and insufficient transmission efficiency. Summary of the Invention

[0004] In view of this, the present disclosure provides a method, device, and storage medium for data management and distribution in a machine learning platform.

[0005] This method manages and shares data based on the Git version control system. When facing the requirements of multi-availability zone training tasks, a two-level storage structure is set up in each availability zone. The primary storage is used to mirror the git source station, and the secondary storage is used to store files according to the commit id. The Argo Workflow is used for task DAG scheduling. By sharing the storage space and distribution operations among multiple tasks, the efficiency of data storage and transmission can be effectively improved.

[0006] According to one aspect of the present application, a method for data management and distribution in a machine learning platform is provided, including:

[0007] Managing and sharing data based on the Git version control data system. The data storage includes multiple availability zones of machine learning training clusters, and each of the availability zones is provided with a two-level storage structure, where the primary storage is used to mirror the git source station;

[0008] The secondary storage is used to store files according to the commit id, and each data version is stored in the secondary storage, and each version has a corresponding commit id identifier. Different training tasks share and call the secondary storage.

[0009] In a possible implementation, the Argo Workflow is used for task DAG scheduling. When data needs to be called during a machine learning training, the to-be-called commit id identifier of the data to be called is determined;

[0010] Determine whether the data to be called is ready according to the commit id identifier to be called through the custom resources of Kubernetes;

[0011] When the data to be called is ready, the currently executing machine learning training uses the data to be called for training.

[0012] In a possible implementation, when determining whether the data to be called is ready according to the commit id identifier to be called through the custom resources of Kubernetes, it further includes:

[0013] When the data to be called is not ready, check whether the CR corresponding to the commit id identifier to be called is in execution according to the commit id identifier to be called, and obtain a judgment result;

[0014] When the CR is in execution, suspend the CR to be executed;

[0015] When the CR to be executed can be executed, the data to be called is ready.

[0016] In a possible implementation, the judgment result obtained by checking whether the CR corresponding to the commit id identifier to be called is in execution further includes:

[0017] When the CR is not in execution, submit the CR application for CR data preparation.

[0018] In a possible implementation, the CR data preparation includes determining whether the specified commit id is obtained from the primary storage;

[0019] When the specified commit id is obtained from the primary storage, based on the obtained specified commit id, use the data corresponding to the obtained commit id as the secondary storage;

[0020] When the specified commit id is not obtained, trigger the fetch operation of the primary storage to obtain the specified commit id again.

[0021] In a possible implementation, when the specified commit id is not obtained, triggering the fetch operation of the primary storage includes:

[0022] When the specified commit id is not obtained, before triggering the fetch operation of the primary storage, check whether there is a fetch operation of the primary storage in execution;

[0023] When there is a fetch operation in the first-level storage during execution, the CR enters the waiting mode;

[0024] When the fetch operation of the first-level storage is completed, the CR that entered the waiting mode fetches the specified commit id again.

[0025] In a possible implementation, when data needs to be called during the machine learning training, Argo Workflow is used for task scheduling, and a training task is described as a DAG. When the DAG executes to the data initialization node, it checks whether the data preparation is completed;

[0026] Among them, the data initialization node of the DAG is implemented based on the CR.

[0027] In a possible implementation, the CR readiness status check is implemented through a custom resource controller;

[0028] Among them, the status of the CR includes at least one of an initial state, first-level data preparation in progress, first-level data preparation completed, data preparation failed, and data preparation in progress.

[0029] According to another aspect of the present application, there is provided a device for data management and distribution of a machine learning platform, including: a processor; a memory for storing processor-executable instructions; wherein, the processor is configured to execute the above method.

[0030] According to another aspect of the present application, there is provided a non-volatile computer-readable storage medium, on which computer program instructions are stored, wherein the computer program instructions, when executed by a processor, implement the above method.

[0031] According to the following detailed description of the exemplary embodiments with reference to the accompanying drawings, other features and aspects of the present disclosure will become clear. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] The accompanying drawings, which are included in and constitute a part of this specification, illustrate exemplary embodiments, features, and aspects of the present disclosure and are used to explain the principles of the present disclosure.

[0033] Figure 1 A block diagram of the secondary storage structure of the training cluster corresponding to the available zone showing an embodiment of the present application;

[0034] Figure 2 A schematic diagram showing the DAG execution process of an embodiment of the present application;

[0035] Figure 3 A schematic diagram showing the specific logic process of the custom controller of an embodiment of the present application;

[0036] Figure 4 A device structure block diagram showing data management and distribution of a machine learning platform according to an embodiment of the present application. Detailed implementation manners

[0037] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. Identical reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.

[0038] The term "exemplary" used herein means "serving as an example, embodiment, or illustration". Any embodiment described herein as "exemplary" is not necessarily to be construed as superior or better than other embodiments.

[0039] In addition, for a better description of the present disclosure, numerous specific details are given in the following detailed implementation manners. Those skilled in the art should understand that the present disclosure can also be implemented without some specific details. In some instances, methods, means, elements, and circuits well known to those skilled in the art are not described in detail so as to highlight the gist of the present disclosure.

[0040] Figure 1 A method for data management and distribution of a machine learning platform according to an embodiment is shown. In the method of this embodiment, data is managed and shared based on a Git version control data system. The data storage includes multiple availability zones for machine learning training clusters, and each of the availability zones is provided with a two-level storage structure. Among them, the primary storage is used to mirror the Git source site, so that in the availability zone corresponding to the training cluster, a complete Git source site mirror is maintained. Specifically, the number of availability zones for data storage corresponds to the number of clusters for machine learning training using the data stored on this storage platform. Of course, in the actual processing process, the specific number of availability zones can be preset, or the corresponding available areas can be automatically divided by the platform when the machine learning is started. Specifically, the corresponding processing method can be executed according to the requirements in the practice process. Moreover, for the specific number of available areas, it is at least more than 2, and can also be more, such as 3, 4, 5,..., or even more than ten or more dozens, hundreds can be achieved. It can also be set according to the operating ability of the data storage management platform.

[0041] A secondary storage structure is set up in each available zone. The primary storage is used to mirror the git source site; the secondary storage is used to store files according to the commit id, and each data version is stored in the secondary storage, and each version has a corresponding commit id. Different training tasks within a single training cluster share the secondary storage in the corresponding available zone of the training cluster. During the call process, the source data is the mirrored data obtained by mirroring the git source site data in the primary storage of the corresponding available zone of the training cluster; the secondary storage stores the version data distributed based on the version identification of the mirrored data in the primary storage.

[0042] In the solution of this embodiment, by setting up the secondary storage structure to store the version data called by the training cluster, different training tasks within the training cluster can call the backup of the data version that has been called by other tasks, so that multiple training tasks within the training cluster can call and share the version data with each other, reducing the data storage volume, and also improving the speed of data preparation and call, thereby improving the machine learning and training efficiency.

[0043] Specifically, the git source site is mirrored in the primary storage of the corresponding available zone of the training cluster. The git fetch operation is periodically executed to achieve the synchronization between the primary storage and the git source site. In a specific embodiment of this application, a timing task is set through the crontab instruction to perform the git fetch operation at 2 o'clock in the morning every day to achieve the synchronization between the primary storage and the git source site.

[0044] Among them, the running cycle of git fetch can be adjusted according to the actual task requirements and is not specifically limited in this application.

[0045] The entire storage platform is divided into a primary available zone and the available zones corresponding to multiple specific training clusters. A primary available zone can be set on the data platform, and the git source site data is stored in this zone. For multiple training clusters, there are multiple corresponding available zones. Each available zone is used to store the mirror repository of the git source site and the data called during the training of the corresponding training cluster tasks. Specifically, in an embodiment of this application, in the available zone A corresponding to training cluster A, the primary storage takes the gitrepo path as the unit to perform snapshot caching on the git source data; the secondary storage takes the git commit id as the unit to perform snapshot caching on the git data. For a primary available zone, multiple training clusters and the corresponding available zones of these training clusters can exist simultaneously; when more than one training task in a training cluster uses the data of the same version, based on the specified commitid, the version data in the secondary storage can be called by multiple training tasks, thereby realizing the distribution and sharing of version data among multiple tasks.

[0046] This application uses Argo Workflow for training task scheduling. For a training task in a training cluster, different execution stages such as data preparation and training execution of a single training task are described as a directed acyclic graph (DAG, Directed Acyclic Graph). The steps of the task are described as nodes, and the dependencies between tasks are described as directed edges; the dependency relationship is realized by specifying the output of one node as the input of another node. In this DAG, each node has specific execution logic and execution conditions. A dedicated data initialization node is set at the starting point of the DAG for the preparation of training data.

[0047] Among them, the data initialization node is defined through the custom resource (CR, Custom Resource) of Kubernetes. Each CR uses two fields, git repo path and commit id, to uniquely identify the data of a specific version. Among them, the git repo path field represents the primary storage of specific data, and the commit id represents the secondary storage of specific data. For tasks with the same git repo path field but different commit ids, only the data path of the primary storage is shared, and the version data of the secondary storage is not shared. Perform a checkout operation in the primary storage to retrieve the specified commit id, and obtain the corresponding version data from the primary storage image based on the retrieved commit id, and use the version data corresponding to the obtained commit id as the secondary storage. When multiple training tasks use the same CR, as long as the CR has been completed, this data can be reused, thereby achieving the purpose of sharing the CR node in multiple DAGs, and then continuing with the subsequent processes of the DAG.

[0048] Specifically, the process of Argo Workflow executing the DAG is as Figure 2 shown. When the DAG starts to execute until the data initial node CR, the Kubernetes API system first checks whether the CR is ready (i.e., the data preparation is completed). If the CR already exists and the status is completed, the system directly executes the downstream nodes of the training task; if it is not ready, wait for the data pull to complete; if the CR is not yet ready, the data pull is not completed, or it does not exist (not used by other tasks). Start the initialization node, continue with the data pull and preparation work, and re-check the CR status. The CR status is mainly divided into the initial state (Declared) where the data is declared, data preparation in progress (Fetching), data preparation completed (FetchchCompleted), data preparation failed (Failed), and data preparation completed (Ready).

[0049] When a CR is created, it indicates that a training task declares its need for a specific version of data. For a specific version of data corresponding to a CR, it can be used normally only when the status of the CR is Ready. The complete specific version of data to be called is placed in the secondary storage within the available area. Subsequently, if there are other training tasks pulling the same version of data, they can directly call it from the secondary storage without having to perform repeated data pulls.

[0050] Furthermore, based on the Kubernetes API extension mechanism, a custom controller is implemented to process the CR, prepare the specific version of data declared in the CR, and feedback the status changes during the processing to the CR so that other services on Kubernetes can also obtain the CR status.

[0051] Specifically, as Figure 3 shown, the specific logic of the custom controller when checking whether the CR is ready is as follows: When there is version data to be called, a CR is specified according to the git repo path and commit id provided by the training task, and it is judged whether the CR is in the Declared state; if the CR is in the Declared state, a git checkout operation is performed on the primary storage; if the CR is not in the Declared state, the Declared state of the CR is first updated to True, and then a git checkout operation is performed on the primary storage.

[0052] Then it is checked whether the above-executed checkout operation is successful. If successful, the Ready state of the CR is updated to True, the Fetching state is updated to False, and the CR data preparation is completed;

[0053] If the checkout operation fails due to unsynchronized data, a fetch operation on the primary storage is triggered to update the git mirror data of the primary storage from the git source. Preferably, before executing the fetch operation, the controller first judges whether the CR is in the Fetching state, that is, whether there is a running fetch operation on the primary storage, and waits for the current fetch operation to complete;

[0054] If the current CR is not in the Fetching state, that is, there is no running fetch operation, then the FetchCompleted state of the CR is checked. When the FetchCompleted state is True, the checkout operation is directly executed; otherwise, the FetchCompleted state is updated to True, and then a fetch operation is performed on the primary storage.

[0055] When the fetch is successful, update the FetchCompleted status of the CR to True and the Fetching status of the CR to False; when the fetch is unsuccessful, update the Failed status of the CR to True and the Fetching status of the CR to False;

[0056] During the above CR preparation process, there may be other training tasks that also depend on the CR with the same commit id. These tasks also enter the suspended state and wait for the CR to be prepared. When the status of the CR changes to completed, based on the task activation mechanism of Argo Workflow, the system automatically activates the suspended tasks and continues with the DAG scheduling and execution.

[0057] Reuse the data corresponding to the CR for the tasks in the suspended state that subscribe to the CR in the order of the suspended queue.

[0058] When the CR data preparation is completed in the above process, the training task calls the version data of the CR identified by the specified git repo path and commit id to complete the task training.

[0059] Provide the expected status of the CR to the system, that is, the specified commit id, in the form of a YAML configuration file, and match the actual status of the CR with the expected status to make the current actual status of the CR reach the expected status.

[0060] Among them, the process of the actual status of the CR approaching the expected status is a conventional technical process based on the Kubernetes platform and will not be elaborated in this application.

[0061] In a possible way,

[0062] When multiple versions of the data are in use at the same time, only perform one cross-cluster data transfer git fetch, and then distribute it into multiple versions through secondary storage within the training cluster to avoid network load caused by multiple cross-cluster transfer operations.

[0063] After secondary storage according to the data version, multiple tasks can share the data of the same version, reducing the work of duplicate data storage.

[0064] In a possible implementation, multiple tasks share the same version of data. Based on the required commit id, it is checked whether the version data corresponding to the required commit id can be directly called. If the required data is being called by other tasks, the current task that needs this version of data is suspended. After the previous task that used this version of data is completed, other tasks that are suspended and subscribed to this version of data are sequentially executed according to the preset execution order, completing data sharing, avoiding multiple loading of the same version of data during the execution of multiple tasks, reducing the work of redundant data storage, and thus improving the efficiency of data storage.

[0065] Furthermore, according to another aspect of the present disclosure, there is also provided a device for data management and distribution of a machine learning platform. Refer to Figure 4 , the device for data management and distribution of the machine learning platform in the embodiments of the present disclosure includes a configuration to implement any of the foregoing methods when executing executable instructions.

[0066] Here, it should be noted that in the device for data management and distribution of the machine learning platform in the embodiments of the present disclosure, it can be connected through a bus or other means, and specific limitations are not provided here.

[0067] The memory 110, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and various modules, such as: the programs or modules corresponding to the method for data management and distribution of the machine learning platform in the embodiments of the present disclosure. The processor 120 executes various functional applications and data processing of the device for data management and distribution of the machine learning platform by running the software programs or modules stored in the memory.

[0068] According to another aspect of the present application, there is also provided a non-volatile computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, any of the foregoing methods is implemented.

[0069] The method for data management and distribution of the machine learning platform according to the above embodiments of the present application can manage and share data through the Git version control system. When facing the training task requirements of multiple availability zones, a two-level storage structure is set up in each availability zone. The primary storage is used to mirror the git source station, and the secondary storage is used to store files according to the commit id. The Argo Workflow is used for task DAG scheduling. When multiple versions of data are in use at the same time, only one cross-cluster data transfer, git fetch, is required, and then the data is distributed into multiple versions through the secondary storage within the training cluster, avoiding the network load caused by multiple cross-cluster transfer operations. After secondary storage according to the data version, multiple tasks can share the data of the same version, reducing the work of duplicate data storage. The present application effectively improves the data storage and transmission efficiency by realizing the sharing of storage space and distribution operations among multiple tasks.

[0070] The embodiments of the present application have been described above. The above description is exemplary and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments. The selection of the terms used herein is intended to best explain the principles of the embodiments, the practical application, or the technical improvement of the technology in the market, or to enable other ordinary skill in the art in the technical field to understand the embodiments disclosed herein.

Claims

1. A data management and distribution method for a machine learning platform, characterized in that: Based on the Git version control data system to manage and share data, data storage includes multiple machine learning training cluster availability zones, each of which is set up with a two-level storage structure, where the first-level storage is used to mirror the git source station; The secondary storage is used to store files by commit id, and each data version is stored in the secondary storage, and each version has a corresponding commit id identifier. Different training tasks share and call the secondary storage; Using Argo Workflows to schedule training tasks includes: When data needs to be called in a machine learning training, the commit id of the data to be called is determined; According to the commit id to be called, determine whether the data to be called is ready through the custom resources of Kubernetes; When the data to be called is prepared, the currently executed machine learning training uses the data to be called for training; The step of determining whether the data to be called is ready through the custom resource of Kubernetes according to the commit id to be called also includes: When the data to be called is not ready, check whether the CR corresponding to the commit id to be called is being executed according to the commit id to be called, and obtain a judgment result; When the CR is being executed, suspend the CR to be executed; When the CR to be executed can be executed, the data to be called is ready.

2. The data management and distribution method of the machine learning platform according to claim 1, characterized in that: The judgment result of checking whether the CR corresponding to the commit id identifier to be called is obtained during execution also includes: When the CR is not being executed, the CR application is submitted and CR data preparation is performed.

3. The data management and distribution method of the machine learning platform according to claim 2, characterized in that: The CR data preparation includes: Determine whether the specified commit id is obtained from the primary storage; When a specified commit ID is obtained from the primary storage, based on the obtained specified commit ID, data corresponding to the obtained commit ID is used as the secondary storage; When the specified commit id is not obtained, a fetch operation of the primary storage is triggered to obtain the specified commit id again.

4. The data management and distribution method of the machine learning platform according to claim 3, characterized in that: When the specified commit id is not obtained, the fetch operation of the primary storage is triggered, including: When the specified commit id is not obtained, before triggering the fetch operation of the primary storage, check whether there is a fetch operation of the primary storage being executed; When there is a fetch operation of the primary storage being executed, the CR enters a waiting mode; When the fetch operation of the primary storage is completed, the CR that enters the waiting mode obtains the specified commit id again.

5. The data management and distribution method of the machine learning platform according to any one of claims 1 to 4, characterized in that: When data needs to be called during the machine learning training, Argo Workflow is used to schedule tasks and describe a training task as a DAG. When the DAG is executed to the data initialization node, check whether data preparation is completed. The data initialization node of the DAG is implemented based on CR.

6. The data management and distribution method of the machine learning platform according to any one of claims 1 to 4, characterized in that: The CR preparation completion status check is implemented through a custom resource controller; The status of the CR includes at least one of an initial status, first-level data preparation in progress, first-level data preparation completed, data preparation failed, and data preparation completed.

7. A data management and distribution device for a machine learning platform, characterized in that: include: processor; a memory for storing processor-executable instructions; Wherein, the processor is configured to implement the method described in any one of claims 1 to 6 when executing the executable instructions.

8. A non-volatile computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Machine learning method and device based on lake and cabin integration

    CN117787432A