Pre-loading method and device of training data, computer equipment and storage medium

By obtaining cluster node information, determining cache priority, and preloading training data when a cache node fails, the access delay problem caused by cache node failure is solved, and the efficiency and stability of model training are improved.

CN120780360AActive Publication Date: 2025-10-14CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511285708.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-10
Publication Date
2025-10-14
Estimated Expiration
2045-09-10

AI Technical Summary

Technical Problem

During the model training process, cache node failures result in high latency and slow reading speeds for the first read access of training data, affecting the recovery efficiency of training tasks. Cache node failures in existing technologies are unpredictable, and cache nodes that newly undertake tasks fail to load the required training data in advance.

Method used

When a current cache node failure is detected, the node information of multiple available cache nodes in the target cluster is obtained, the cache priority is determined, the candidate available cache nodes are screened out, and the target available cache nodes are controlled to preload training data. The data access path is optimized through intelligent node screening and priority-driven preloading strategies.

Benefits of technology

Through intelligent node screening and priority-driven preloading strategies, data reading latency is reduced, the waiting time for training tasks to load data is shortened, and fault tolerance, resource utilization, and business stability are improved, especially significantly improving efficiency in large-scale training scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120780360A_ABST
    Figure CN120780360A_ABST
Patent Text Reader

Abstract

The invention relates to a training data preloading method and device, computer equipment and a storage medium. Comprising the following steps: acquiring node information of a plurality of available cache nodes in a target cluster and training data of a model training task under the condition of detecting that a current cache node has a fault; according to the node information of the plurality of available cache nodes, cache priorities of the plurality of available cache nodes for loading training data are determined; determining at least one candidate available cache node from the plurality of available cache nodes according to the training data; and determining a target available cache node from the at least one candidate available cache node according to the cache priority, and controlling the target available cache node to pre-load the training data. Therefore, the fault-tolerant capability, the resource utilization rate and the service stability of the target cluster can be remarkably improved in scenes with high data timeliness requirements such as large-scale training.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the field of model training, and particularly relates to a preloading method and device of training data, a computer device and a storage medium. BACKGROUND

[0002] In actual application, part of data in the remote storage can be cached into the local cache node through artificial pre-configuration, so as to improve the read-write speed of the data.

[0003] In the related art, the model training process usually takes a long time, and the cache node may fail during the process. In order to ensure the continuous operation of the model training task, the model training task is usually rescheduled to other cache nodes to resume the training after the cache node fails. However, the cache node failure is unpredictable, and the cache node that takes over the task usually fails to load the required training data in advance, so when the cache node reads the training data for the first time, it still faces the problem of high access delay and slow reading speed, which affects the recovery efficiency of the training task. SUMMARY

[0004] Therefore, the present disclosure provides a preloading method and device of training data, a computer device and a storage medium to solve the problems in the related art.

[0005] In a first aspect, the present disclosure provides a preloading method of training data, which comprises: in the case that a current cache node is detected to fail, acquiring node information of a plurality of available cache nodes in a target cluster and training data included in the current cache node; respectively determining cache priorities of the plurality of available cache nodes for loading the training data according to the node information of the plurality of available cache nodes; determining at least one candidate available cache node from the plurality of available cache nodes according to the training data; determining a target available cache node from the at least one candidate available cache node according to the cache priorities, and controlling the target available cache node to pre-load the training data.

[0006] In a second aspect, the present disclosure provides a preloading device of training data, which is applied to the preloading method of training data as in the first aspect, and comprises: an acquisition module, configured to acquire node information of a plurality of available cache nodes in a target cluster and training data included in a current cache node in the case that the current cache node is detected to fail; a first determination module, configured to respectively determine cache priorities of the plurality of available cache nodes for loading the training data according to the node information of the plurality of available cache nodes; a second determination module, configured to determine at least one candidate available cache node from the plurality of available cache nodes according to the training data; and a control module, configured to determine a target available cache node from the at least one candidate available cache node according to the cache priorities, and control the target available cache node to pre-load the training data.

[0007] In a third aspect, the present disclosure provides a computer device, comprising a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the preloading method of training data.

[0008] In a fourth aspect, the present disclosure provides a computer readable storage medium, having stored thereon a computer program / instruction, which, when executed by a processor, implements the steps of the preloading method of training data.

[0009] According to a fifth aspect, the present disclosure provides a computer program product, which, when executed by a processor, implements the steps of the preloading method of training data.

[0010] The above at least one technical solution adopted by the embodiments of the present disclosure can achieve the following beneficial effects: by detecting that the current cache node has failed, the node information of a plurality of available cache nodes in the target cluster and the training data of the model training task are obtained; the cache priority of the plurality of available cache nodes to load the training data is determined according to the node information of the plurality of available cache nodes respectively; at least one candidate available cache node is determined from the plurality of available cache nodes according to the training data; the target available cache node is determined from the at least one candidate available cache node according to the cache priority, and the target available cache node is controlled to pre-load the training data.

[0011] As can be seen, the embodiments of the present disclosure can immediately trigger a multi-stage processing flow when it is detected that the current cache node has failed. First, the node information of the available cache nodes in the target cluster can be obtained. Then, the cache priority of each available cache node is determined through the node information. Then, the candidate available cache nodes with cache potential are screened based on the characteristics of the training data. Finally, the optimal target available cache node is selected according to the cache priority to perform the preloading operation of the training data. In this way, the intelligent node screening method can be used to replace manual intervention to avoid cache resource overload. At the same time, the preloading strategy driven by the priority optimizes the data access path of the failed current cache node from remote reading to local cache reading, reduces the delay rate of data reading, and improves the efficiency. Moreover, the design of preloading completes the deployment of training data in advance, greatly shortens the time for the training task to wait for data loading, reduces the empty window period of the training process, and significantly improves the fault tolerance, resource utilization, and business stability of the target cluster, especially in large-scale training and other scenarios with high requirements for data timeliness. BRIEF DESCRIPTION OF DRAWINGS

[0012] The above and other purposes, features, and advantages of the present disclosure will become more apparent through a more detailed description of the embodiments of the present disclosure in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the present disclosure and are not intended to limit the present disclosure. In the drawings, the same reference numerals generally represent the same components or steps.

[0013] Figure 1 A schematic diagram of the structure of a cache coordination system provided by an exemplary embodiment of the present disclosure; Figure 2 A flowchart of a method for preloading training data provided by an exemplary embodiment of the present disclosure; Figure 3 A schematic structural diagram of a training data preloading device provided in one embodiment of the present disclosure; Figure 4 A schematic structural diagram of an electronic device provided in one embodiment of the present disclosure; Figure 5 A schematic diagram of the structure of a computer system provided in one embodiment of the present disclosure; Figure 6 A schematic diagram of a computer program product provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0014] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0015] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0016] The term "including" and its variations used in this document are open inclusions, that is, "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one other embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the description below. It should be noted that the concepts of "first", "second", etc. mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0017] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0018] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0019] With the rapid development of artificial intelligence technology, the demand for model training in cloud environments continues to grow, especially in large-scale model training scenarios, which are characterized by high training costs, large computing node scale, and long training cycles. Typically, during model training, training data is frequently read, and during this process, GPU nodes are often idle and waiting, resulting in reduced computing resource utilization and significant resource waste. A GPU node, short for Graphics Processing Unit Node, is a hardware unit that integrates graphics processor hardware and has independent computing, storage, and network communication capabilities. It is a core component for building GPU node clusters or single-machine high-performance computing systems.

[0020] Related technologies offer low-cost cloud storage services, such as object storage. Using object storage as the underlying storage architecture during model training can help further control costs. However, object storage is typically located in a different physical or logical network area from the model training compute cluster. This means that GPU nodes must access data directly from object storage remotely over the network. This not only introduces high network transmission overhead but also results in significant data read latency, becoming a bottleneck for training performance.

[0021] To address the above issues, related technologies use a distributed caching method, which caches part of the data in the remote storage to the local cache node, so that the GPU node can directly read the training data from the local cache node, thereby converting remote reads into local reads and writes, significantly improving I / O performance. However, for the distributed caching method, if the required training data has not been cached, it is still necessary to read the data from the remote storage and then store it in the cache. In this case, the reading speed of the training data is even lower than directly accessing the remote storage. To address the problem of poor first-time read performance, the required data set can be pre-loaded into the distributed cache system before the training task is executed through manual pre-configuration to ensure that the data can be obtained from the local cache on the first read, avoiding remote access delays.

[0022] However, manual pre-configuration still has its shortcomings, especially in multi-task parallel and fault-tolerant scenarios, where its flexibility and degree of automation are limited. First, in model training scenarios, multiple model training tasks often run concurrently in the same cluster, and each model training task only occupies part of the cache nodes in the cluster. If all the training data required for all model training tasks are pre-loaded into the distributed cache system manually, it may not be able to accommodate all the training data due to limited cache capacity, resulting in cache resource competition and reduced efficiency. Secondly, the model training process usually takes a long time, during which cache nodes may fail. In order to ensure the continuous operation of model training tasks, model training tasks are usually rescheduled to other cache nodes to resume training after cache node anomalies. However, cache node failures are unpredictable, and cache nodes that newly undertake tasks often fail to load the required training data in advance. Therefore, when the cache node reads the training data for the first time, it still faces high access latency and slow reading speed, affecting the recovery efficiency of the training task.

[0023] In response to the above problems, an embodiment of the present disclosure provides a method for preloading training data, including obtaining node information of multiple available cache nodes in a target cluster and training data of a model training task when a current cache node is detected to have failed; determining cache priorities of multiple available cache nodes for loading training data based on the node information of the multiple available cache nodes; determining at least one candidate available cache node from the multiple available cache nodes based on the training data; determining a target available cache node from the at least one candidate available cache node based on the cache priority, and controlling the target available cache node to preload training data.

[0024] It can be seen that the embodiment of the present disclosure can immediately trigger a multi-stage processing flow when a fault is detected in the current cache node. First, the node information of the available cache nodes in the target cluster can be obtained; then the cache priority of each available cache node can be determined through the node information; then the candidate available cache nodes with cache potential are screened out based on the training data characteristics; finally, the optimal target available cache node is selected according to the cache priority to perform the preloading operation of the training data. In this way, an intelligent node screening method can be used to replace manual intervention to avoid cache resource overload; at the same time, the data access path of the faulty current cache node is optimized from remote reading to local cache reading through the priority-driven preloading strategy, which reduces the data reading delay rate and improves; and, the preloading design completes the training data deployment in advance, greatly shortens the time for training tasks to wait for data loading, and reduces the window period of the training process. Especially in scenarios with high requirements for data timeliness such as large-scale training, it can significantly improve the fault tolerance, resource utilization and business stability of the target cluster.

[0025] The training data preloading method provided by the embodiments of the present disclosure may be executed by a terminal or by a chip applied to the terminal.

[0026] Exemplarily, the above-mentioned terminals may include one or more of mobile phones, tablet computers, wearable devices, vehicle-mounted devices, laptop computers, ultra-mobile personal computers (UMPCs), netbooks, personal digital assistants (PDAs), and wearable devices based on augmented reality (AR) and / or virtual reality (VR) technologies, and may also include but not be limited to smart terminals such as remote control devices, wearable devices, street lights, and home appliances. The embodiments of the present disclosure do not impose specific restrictions on this.

[0027] Figure 1 This is a schematic diagram of a cache coordination system provided by an exemplary embodiment of the present disclosure. Figure 1 Shown, including: The intelligent computing platform 101, as the upper-level coordination module, is used to manage the scheduling and resource coordination of model training tasks and is the brain of the entire process; the cache coordinator 102 is used to undertake the functions of cache strategy formulation, cache node status management, and coordination of cache resource allocation; the training node 103 includes the cache node 1031, which is a temporary warehouse for training data and is used to store data; the object storage 104: provides data support for the training node.

[0028] like Figure 1As shown, when the intelligent computing platform 101 sends the model training task to each training node 103, the cache coordinator 102 can communicate with the intelligent computing platform 101 and the training node 103 in a two-way manner to dynamically allocate cache resources; the training node 103 reads and writes training data from the object storage 104 on demand, and the cache node 1031 uses the cache strategy to accelerate the access to the training data, ensure the efficient execution of the model training task, and realize the closed-loop process of task scheduling, cache coordination and data access, so as to adapt to the requirements of large-scale distributed model training for high concurrency and low-latency access to data.

[0029] Figure 2 This is a flow chart of a method for preloading training data provided by an exemplary embodiment of the present disclosure. Figure 2 As shown, the embodiment of the present disclosure can be applied to Figure 1 The cache coordinator shown in the figure has the following specific preloading methods for training data: S201: Upon detecting a failure of a current cache node, obtain node information of multiple available cache nodes in a target cluster and the training data included in the current cache node. It should be understood that the training data herein refers to the training data required for executing a model training task, and one piece of training data may include one or more training sub-data corresponding to one model training task.

[0030] In some embodiments, the node information may include at least one of the remaining cache space of an available cache node, cached data, the amount of cached data, the number of failures, and the time of the last abnormality; and the training data may be the data corresponding to the model training task being executed by the current cache node that failed. Based on this, after the cache coordinator receives the failure signal of the current cache node, it can immediately obtain the node information of all available cache nodes in the target cluster, and then use one of the multiple available cache nodes to pull the training data. In other words, the available cache nodes can be used to directly take over the data needs of the failed cache node, thereby minimizing the training interruption period and improving the efficiency of model training.

[0031] S202 : Determine cache priorities of multiple available cache nodes for loading training data according to node information of multiple available cache nodes.

[0032] In some embodiments, the cache priority of any available cache node for loading training data may be determined based on the remaining cache space, cached data, amount of cached data, number of failures, and last abnormality time of the available cache node.

[0033] Specifically, the cache priority of the i-th available cache node when loading the j-th training data can be defined as , here, The calculation formula is as follows:

[0034] in, for and The result of the sum after the sigmoid function transformation can ensure that The value of is in the range of (0, 1), which makes it easy to set the priority threshold. Node information includes one of the following: available cache node remaining cache space, cached data, cached data volume, number of failures, and the time of the last failure, with k = 1, 2, 3, or 4. The Sigmoid function is a common S-shaped activation function widely used in machine learning (especially logistic regression and neural networks). Its core function is to map any real-number input value to the [0, 1] interval, achieving "numerical normalization" or "probabilistic output."

[0035] And for , can be defined separately, the specific calculation is as follows:

[0036] in, Indicates the influence of the remaining cache space of the available cache node on the cache priority. Here, the more remaining cache space there is, the greater the possibility that the available cache node is used. is the scale factor, which can be used as a custom The proportion of cache priority; The remaining cache space available for the available cache nodes; The total cache space size cached for available cache nodes.

[0037]

[0038] in, Indicates the impact of the amount of cached data of available cache nodes on cache priority; is the scale factor, which can be used as a custom The proportion of cache priority; Indicates the number of model training tasks that have been cached on the available cache nodes. Here, if the available cache node has already cached a large number of training tasks, you can find other available cache nodes instead of loading all training datasets into the same available cache node; The importance of the training data is scored. Here, if the importance row score of a certain cached data in an available cache node is very high, the influence of other cached data can be ignored, and the cached data with a relatively high importance row score is still loaded on the available cache node.

[0039]

[0040] wherein, represents the influence of the last abnormal time of the available cache node on the priority; is a proportional coefficient, which can be used as a custom proportion in the cache priority; is the difference between the last failure time of the available cache node and the current time.

[0041]

[0042] wherein, represents the influence of the number of failures of the available cache node on the priority; is a proportional coefficient, which can be used as a custom proportion in the cache priority, is the number of cumulative failures of the available cache node.

[0043] Based on this, the priority calculation mechanism of the embodiments of the present disclosure fuses the four core indicators of the remaining cache space, the data amount of the cached data, the last failure time and the number of failures of the available cache node, and forms a priority value in the interval (0, 1) after sigmoid function normalization processing, which not only avoids cache overflow by preferentially selecting the available cache node with sufficient remaining cache space, but also prevents the training data from being excessively concentrated on a single available cache node by , and tends to select the stable available cache node with long failure interval and few failures, and at the same time, allows the proportional coefficient to flexibly adjust the weight of each factor to adapt to different scene requirements, and finally realizes the scientific quantification of the priority of the available cache node to load the training data, provides a precise decision basis for the efficient migration of the training data after failure, guarantees the reliability and efficiency of data loading, and also takes into account the balanced utilization of cluster resources. S203, determining at least one candidate available cache node from the plurality of available cache nodes according to the training data.

[0044] In some embodiments, the priority of all available cache nodes can be obtained after which a priority threshold can be set to filter out

[0045] the available cache node with a priority lower than the priority threshold. ​Available cache nodes with a value greater than this threshold are selected as candidate available cache nodes. If the filtering result is empty, the top three nodes with the highest priority are selected as candidate available cache nodes. Alternatively, the cache priority corresponding to the available cache nodes that have already run training data can be set to 0, and then the available cache nodes that do not contain training data can be filtered out from the remaining available cache nodes as candidate available cache nodes. Based on this, based on the quantitative cache priority indicator, candidate available cache nodes with high adaptability can be quickly identified, avoiding the waste of resources of candidate available cache nodes with low cache priority.

[0046] S204: Determine a target available cache node from at least one candidate available cache node according to the cache priority, and control the target available cache node to preload training data.

[0047] In some embodiments, when there is one candidate available cache node, the candidate available cache node can be determined as the target available cache node; when there are multiple candidate available cache nodes, the candidate available cache node with the highest cache priority can be determined as the target available cache node based on the cache priority of at least one candidate available cache node.

[0048] Based on this, the disclosed embodiment can immediately trigger a multi-stage processing flow upon detecting a failure in the current cache node. First, the node information of the available cache nodes in the target cluster can be obtained; then, the cache priority of each available cache node can be determined through the node information; then, candidate available cache nodes with cache potential can be screened out based on the training data characteristics; finally, the optimal target available cache node can be selected based on the cache priority to perform the preloading operation of the training data. In this way, an intelligent node screening method can be used to replace manual intervention to avoid cache resource overload; at the same time, the data access path of the failed current cache node is optimized from remote reading to local cache reading through a priority-driven preloading strategy, which reduces the data reading delay rate and improves; moreover, the preloading design completes the training data deployment in advance, greatly shortens the time for training tasks to wait for data loading, and reduces the window period of the training process. Especially in scenarios with high requirements for data timeliness such as large-scale training, it can significantly improve the fault tolerance, resource utilization and business stability of the target cluster.

[0049] In some embodiments, the training data includes multiple data, and the preloading method of the training data also includes: determining the loading priority of the multiple training data according to the importance scores of the multiple training data respectively; controlling the target available cache nodes corresponding to the multiple training data in sequence, and preloading the multiple training data according to the loading priority order of the multiple training data.

[0050] Specifically, the plurality of training data correspond to a plurality of model training tasks, and the importance scores of the training data can be scored by relevant personnel according to the importance of the model training tasks. For example, assuming that the plurality of training data includes first training data , second training data and third training data At this time, the importance score ranking can be According to the ranking of importance, the importance score of the second training data is the highest, at this time, the second training data can be processed first, that is, the cache priority of the second training data loaded by the plurality of available cache nodes can be determined first, then at least one candidate available cache node is determined from the plurality of available cache nodes according to the second training data, thereby obtaining the candidate available cache node matrix corresponding to the second training data, then the candidate available cache node with the highest cache priority is determined as the target available cache node according to the cache priority corresponding to the plurality of candidate available cache nodes. The specific candidate available cache node matrix is as follows:

[0051] Among them, represents the cache priority of the first candidate available cache node for the second training data; represents the cache priority of the second candidate available cache node for the second training data; represents the cache priority of the i-th candidate available cache node for the second training data.

[0052] When there are a plurality of training data, the target available cache node corresponding to each training data can be determined respectively by referring to the above method, and the plurality of training data are loaded and processed in turn according to the loading priority order corresponding to the plurality of training data. In this way, by determining the loading priority according to the importance score of the training data first, and then controlling the corresponding target available cache node to be preloaded in this order, not only the core training data is ensured to obtain cache resources first, reducing the risk of training blocking caused by loading delay, but also the network bandwidth competition and node resource conflict caused by parallel transmission of multiple data are avoided through ordered loading, improving the overall data loading efficiency. At the same time, the loading order is dynamically adjusted combined with the importance of the training data, which can maximize the availability of high-value data under limited cache resources, making the fault recovery process more in line with the priority order of actual training requirements, further optimizing the rationality of cache resource allocation and the continuity of training tasks.

[0053] In some embodiments, the node information includes cached data in the available cache nodes, and at least one candidate available cache node is determined from multiple available cache nodes based on the training data, including: when the cached data in the available cache node does not match the training data, determining the available cache node as a candidate available cache node.

[0054] Specifically, by screening available cache nodes whose cached data does not match the training data as candidate available cache nodes, it is possible to effectively avoid repeatedly loading training data into available cache nodes that already have the same data. This not only reduces the storage space occupied by redundant caches and the network transmission overhead of data replication, but also improves the utilization efficiency of cache resources by avoiding duplicate storage. At the same time, this screening logic ensures the distributed storage of training data in the cluster, reduces the risk of critical data loss due to a single node failure, makes the cache system more reasonable and reliable in data distribution, and provides better infrastructure support for data access for subsequent training tasks.

[0055] In some embodiments, the node information includes cached data and remaining cache space in the available cache node, and the method for preloading training data also includes: for any candidate available cache node, when the remaining cache space of the candidate available cache node is less than the data volume of the training data, determining the cache data to be removed from the cached data in the candidate available cache node according to the data volume of the training data; removing the cache data to be removed, and re-determining the cache priority of the candidate available cache node when loading the training data.

[0056] Specifically, assume that the remaining cache space of the candidate available cache node is 50GB, and the amount of training data to be loaded is 80GB, with a difference of 30GB. At this point, the 100GB of data cached by the candidate available cache node can be analyzed. For example, the candidate available cache node includes 40GB of cached data for model A training task, 30GB of cached data for model B training task, and 30GB of cached data for model C training task. At this point, the cached data of model B training task with the lowest access frequency can be removed based on rules such as data access frequency and execution priority of model training tasks. After removal, the remaining cache space of the node becomes 80GB, which meets the training data loading requirements. The cache priority of the candidate available cache node is then recalculated. Due to the change in the amount of cached data, the corresponding formula above is The value also needs to be adjusted accordingly.

[0057] It can be seen that the embodiment of the present disclosure can release space by dynamically cleaning up low-priority cache data, thereby solving the problem of insufficient remaining capacity of candidate available cache nodes, ensuring that training data can be loaded smoothly, and improving the elastic utilization of cache resources; at the same time, the removal operation is accurately executed in combination with data characteristics, reducing the impact on high-value cached data, and recalculating the priority to ensure the accuracy of decision-making after the node status changes, which not only avoids the failure of candidate available cache nodes due to space limitations, but also balances the cache requirements of new and old data through refined space management, thereby enhancing the system's adaptability to dynamic storage pressure.

[0058] In some embodiments, the cached data includes data corresponding to multiple cache tasks, and the cached data to be removed is determined from the cached data in the candidate available cache nodes based on the data volume of the training data, including: determining the replacement cost of the data corresponding to each cache task when it is removed based on the importance score and data volume corresponding to each cache task; determining the removal priority of each cache task based on the replacement cost; and determining at least one cache task to be removed from multiple cache tasks based on the data volume and removal priority of each cache task, wherein the data corresponding to at least one cache task to be removed is the cache data to be removed. Here, the replacement cost may include at least one of the time cost, economic cost when removing the data of the cache task, and the time cost, economic cost, etc. consumed in loading the training data.

[0059] Specifically, the replacement cost of cache task t on candidate available cache node y can be calculated according to the following formula: , the specific calculation formula is as follows:

[0060] in, is the size of the data set corresponding to cache task t that has been cached on the candidate available cache node y; is the remaining cache space size that the candidate available cache node y can use to cache data; is the importance score of cache task t cached by candidate available cache node y; r is the cache replacement cost coefficient; R is a value that varies depending on the underlying storage system. Here, candidate available cache node y belongs to available cache node i.

[0061] Based on this, by constructing a multi-dimensional decision chain based on importance score, data volume, replacement cost, and removal priority, refined management of cache data removal can be achieved. Its core value lies in breaking through the traditional cache elimination logic based solely on data volume or access frequency. By quantifying the importance of cache tasks as replacement costs, it avoids the waste of resources that may result from blindly deleting large-capacity, low-value data, and prevents the inefficient occupation of cache space caused by retaining low-priority, small-capacity data. This hierarchical decision-making mechanism can prioritize the retention of high-importance data within limited cache node resources. At the same time, through a comprehensive evaluation of data volume and priority, it ensures that the removal operation can not only meet the new data storage needs, but also minimize the overall cache replacement cost, ultimately improving the resource utilization and data service efficiency of the cache system.

[0062] In some embodiments, a replacement cost for removing at least one cached data item to be removed is determined, and a corresponding use value after loading the training data is determined. If the replacement cost of at least one cached data item to be removed is less than the use value, the cached data to be removed is removed. Alternatively, if the replacement cost of at least one cached data item to be removed is greater than or equal to the use value, a target available cache node is re-determined from at least one candidate available cache node based on the cache priority. Here, the use value may include the economic value generated after the training data is loaded and used.

[0063] Specifically, suppose that the second training data needs to be loaded. At this time, for the candidate available cache node y, if the cache task t corresponds to <M2, then the cache priority can be determined based on the impact of the cached data corresponding to the cache task t in the candidate available cache node y A correction is made, that is, the cached data corresponding to cache task 2 in the candidate available cache node y can be removed, and then the released cache space is used to load the second training data, where M2 represents the usage value of the second training data.

[0064] Based on this, we can After the M2 condition filters out all cache tasks to be removed, the corresponding candidate available cache node y can be recalculated when loading the second training data , where after the remaining cache space of the candidate available cache node y changes, it needs to be modified synchronously and The value of is:

[0065] in, is the remaining cache space size of the candidate available cache nodes after releasing all cache tasks to be removed. is the number of all cache tasks to be removed, Score the importance of the second training data.

[0066] In some embodiments, after removing the cache data to be removed, the cache priority corresponding to the corresponding candidate available cache node can be re-determined. Furthermore, the cache priority of each candidate available cache node whose remaining cache space is less than the data amount of the training data can be recalculated according to the above content. Then, the cache priorities corresponding to all candidate available cache nodes are re-sorted according to the re-determined cache priority to re-determine the target cache node with the highest cache priority as the target cache node for the second training data, and start the process of the second training data loading task.

[0067] In some embodiments, the above method can also be used to calculate only the target cache node, that is, after the target available cache node is determined from at least one of the candidate available cache nodes according to the cache priority, when the remaining cache space of the target available cache node is less than the data volume of the training data, the above method is used to determine the cache data to be removed in the target available cache node, and then the cache data to be removed is moved out, and the training data is loaded and processed using the target available cache node.

[0068] In some embodiments, based on the importance score and data volume corresponding to each cache task, the replacement cost of each cache task's corresponding data when it is removed is determined, and the data loading process is started. When subsequent cache tasks make a second correction to the cache priority, the previously started loading tasks need to be considered. For example, for the second training data corresponding to cache task 2, available cache node 5 can be selected to load the training data. Then, when correcting the cache priority, the cache node 5 can be selected to load the training data. When and Need to be modified to:

[0069] Among them, S2 is the second training data The corresponding data volume.

[0070] In some embodiments, as time goes by, new model training tasks will be continuously undertaken, and the node information of the cache node will change. Therefore, the cache coordinator can be controlled to periodically execute the process of the preloading method of training data in the embodiment of the present disclosure to realize data preloading.

[0071] Based on this, the disclosed embodiment constructs a dynamic and elastic cache management decision logic by introducing a quantitative comparison mechanism of replacement cost and use value, which significantly improves the rationality of cache resource scheduling. When the cost of removing the data to be deleted is lower than the use value of the new data, the removal operation is performed to achieve optimal resource configuration; when the replacement cost is too high or not worth the cost, the priority re-evaluation mechanism of the candidate available cache nodes is activated to avoid irrational cache replacement. This two-tier decision-making model not only prevents the sacrifice of high-cost cache resources for low-value data, but also ensures the flexibility of cache space adjustment through secondary priority screening, and ultimately maximizes the overall utility of the cache system under resource constraints, balancing short-term storage needs and long-term resource utilization efficiency.

[0072] It can be seen that the embodiment of the present disclosure can immediately trigger a multi-stage processing flow when a fault is detected in the current cache node. First, the node information of the available cache nodes in the target cluster can be obtained; then the cache priority of each available cache node can be determined through the node information; then the candidate available cache nodes with cache potential are screened out based on the training data characteristics; finally, the optimal target available cache node is selected according to the cache priority to perform the preloading operation of the training data. In this way, an intelligent node screening method can be used to replace manual intervention to avoid cache resource overload; at the same time, the data access path of the faulty current cache node is optimized from remote reading to local cache reading through the priority-driven preloading strategy, which reduces the data reading delay rate and improves; and, the preloading design completes the training data deployment in advance, greatly shortens the time for training tasks to wait for data loading, and reduces the window period of the training process. Especially in scenarios with high requirements for data timeliness such as large-scale training, it can significantly improve the fault tolerance, resource utilization and business stability of the target cluster.

[0073] The above mainly introduces the solution provided by the embodiment of the present disclosure from the perspective of the server. It can be understood that in order to realize the above functions, the server includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the present disclosure can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present disclosure.

[0074] The embodiments of the present disclosure can divide the server into functional units according to the above method examples. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one management module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiments of the present disclosure is schematic and is only a logical functional division. In actual implementation, there may be other division methods.

[0075] In the case of dividing the functional modules according to the functions, an exemplary embodiment of the present disclosure provides a device for preloading training data, which may be a server or a chip applied to the server. Figure 3 This is a schematic diagram of the structure of a preloading device for training data provided by an embodiment of the present disclosure. Figure 3 As shown, the training data preloading device 300 includes: The acquisition module 301 is configured to acquire node information of multiple available cache nodes in a target cluster and training data included in the current cache node when a fault is detected in the current cache node.

[0076] The first determining module 302 is configured to determine cache priorities of the plurality of available cache nodes for loading the training data according to node information of the plurality of available cache nodes.

[0077] The second determining module 303 is configured to determine at least one candidate available cache node from the plurality of available cache nodes according to the training data.

[0078] The control module 304 is configured to determine a target available cache node from the at least one candidate available cache node according to the cache priority, and control the target available cache node to preload the training data.

[0079] In an optional manner, the training data includes multiple data, and the method further includes: determining the loading priority of the multiple training data according to the importance scores of the multiple training data respectively; controlling the target available cache nodes corresponding to the multiple training data in sequence, and preloading the multiple training data according to the loading priority order of the multiple training data.

[0080] In an optional manner, the node information includes cached data in the available cache node, and determining at least one candidate available cache node from multiple available cache nodes based on the training data includes: when the cached data in the available cache node does not match the training data, determining the available cache node as a candidate available cache node.

[0081] In an optional manner, the node information comprises cached data and remaining cache space in the available cache nodes, and the method further comprises: for any candidate available cache node, when the remaining cache space of the candidate available cache node is less than the data amount of the training data, determining to-be-removed cached data from the cached data in the candidate available cache node according to the data amount of the training data; removing the to-be-removed cached data, and re-determining the cache priority of the candidate available cache node when loading the training data.

[0082] In an optional manner, the cached data comprises data corresponding to a plurality of cache tasks, and the determining to-be-removed cached data from the cached data in the candidate available cache node according to the data amount of the training data comprises: determining a replacement cost of data corresponding to each cache task when the data is removed according to the importance score and the data amount of each cache task; determining a removal priority of each cache task according to the replacement cost of each cache task; and determining at least one to-be-removed cache task from the plurality of cache tasks according to the data amount and the removal priority of each cache task, wherein the data corresponding to the at least one to-be-removed cache task is the to-be-removed cached data.

[0083] In an optional manner, a replacement cost of removing the at least one to-be-removed cached data is determined, and a corresponding use value after loading the training data is determined; in a case where the replacement cost of the at least one to-be-removed cached data is less than the use value, the to-be-removed cached data is removed; or in a case where the replacement cost of the at least one to-be-removed cached data is greater than or equal to the use value, a target available cache node is re-determined from at least one candidate available cache node according to the cache priority.

[0084] The embodiments of the present disclosure further provide an electronic device, comprising: at least one processor; a memory for storing at least one processor-executable instruction; wherein the at least one processor is configured to execute the instruction to implement the steps of the above-mentioned method disclosed by the embodiments of the present disclosure.

[0085] Figure 4 The structural schematic diagram of the electronic device provided by an embodiment of the present disclosure is shown in FIG. 4. As shown in FIG. 4, the electronic device 400 comprises at least one processor 401 and a memory 402 coupled to the processor 401, and the processor 401 can execute the corresponding steps in the above-mentioned method disclosed by the embodiments of the present disclosure. Figure 4

[0086] ​The processor 401 can also be referred to as a central processing unit (CPU), which can be an integrated circuit chip with signal processing capabilities. Each step in the method disclosed in the embodiments of the present disclosure can be performed by hardware integrated logic circuits in the processor 401 or by software instructions. The processor 401 can be a general-purpose processor, a digital signal processor (DSP), an ASIC, a field-programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of the present disclosure can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software modules can be located in memory 402, such as a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The processor 401 reads the information in the memory 402 and, in conjunction with its hardware, completes the steps of the method.

[0087] In addition, when various operations / processes according to the present disclosure are implemented by software and / or firmware, they can be transferred from a storage medium or a network to a computer system having a dedicated hardware structure, for example, Figure 5 The computer system 500 shown is installed with the programs constituting the software. When the various programs are installed, the computer system can perform various functions, including the functions described above. Figure 5 A schematic diagram of the structure of a computer system provided in one embodiment of the present disclosure.

[0088] Computer system 500 is intended to represent various forms of digital electronic computer equipment, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic equipment can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0089] like Figure 5As shown, computer system 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. RAM 503 may also store various programs and data required for the operation of computer system 500. Computing unit 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to bus 504.

[0090] Multiple components within computer system 500 are connected to I / O interface 505, including an input unit 506, an output unit 507, a storage unit 508, and a communication unit 509. Input unit 506 can be any type of device capable of inputting information into computer system 500. Input unit 506 can receive input numeric or character information and generate key input signals related to user settings and / or function control of an electronic device. Output unit 507 can be any type of device capable of presenting information and may include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. Storage unit 508 may include, but is not limited to, a magnetic disk or an optical disk. Communication unit 509 allows computer system 500 to exchange information / data with other devices over a network, such as the Internet, and may include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver and / or chipset, such as a Bluetooth™ device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.

[0091] The computing unit 501 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU node), various specialized artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above. For example, in some embodiments, the methods disclosed in the embodiments of the present disclosure may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed onto the electronic device via the ROM 502 and / or the communication unit 509. In some embodiments, the computing unit 501 may be configured to perform the methods disclosed in the embodiments of the present disclosure by any other suitable means (e.g., via firmware).

[0092] An embodiment of the present disclosure further provides a computer-readable storage medium, wherein when instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the above method disclosed in the embodiment of the present disclosure.

[0093] The computer-readable storage medium in the embodiments of the present disclosure may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus. The computer-readable storage medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the foregoing. More specifically, the computer-readable storage medium may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0094] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0095] Figure 6 Schematic diagram of a computer program product provided by an embodiment of the present disclosure. Figure 6 As shown, the computer program product 600 includes a computer program 601 , wherein the computer program 601 implements the above method disclosed in the embodiment of the present disclosure when executed by a processor.

[0096] In embodiments of the present disclosure, computer program code for performing the operations of the present disclosure may be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer.

[0097] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0098] The modules, components, or units described in the embodiments of the present disclosure may be implemented in software or hardware. The names of the modules, components, or units do not necessarily limit the modules, components, or units themselves.

[0099] The functions described above herein may be at least partially performed by one or more hardware logic components. For example, and without limitation, exemplary hardware logic components that may be used include: a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), an application specific standard product (ASSP), a system on a chip (SOC), a complex programmable logic device (CPLD), and the like.

[0100] The above descriptions are merely some embodiments of the present disclosure and illustrate the underlying technical principles. Those skilled in the art should understand that the scope of the present disclosure is not limited to technical solutions formed by specific combinations of the aforementioned technical features. It also encompasses other technical solutions formed by any combination of the aforementioned technical features or their equivalents, without departing from the aforementioned concepts. For example, a technical solution formed by replacing the aforementioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.

[0101] Although some specific embodiments of the present disclosure have been described in detail by way of examples, those skilled in the art will appreciate that the above examples are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Those skilled in the art will appreciate that modifications may be made to the above embodiments without departing from the scope and spirit of the present disclosure. The scope of the present disclosure is defined by the appended claims.

Claims

1. A method for preloading training data, characterized in that: include: When a fault is detected in the current cache node, obtaining node information of multiple available cache nodes in the target cluster and training data included in the current cache node; Determining cache priorities of the plurality of available cache nodes for loading the training data according to node information of the plurality of available cache nodes respectively; Determine at least one candidate available cache node from the plurality of available cache nodes according to the training data; A target available cache node is determined from the at least one candidate available cache node according to the cache priority, and the target available cache node is controlled to preload the training data.

2. The method according to claim 1, characterized in that The training data includes a plurality of data, and the method further includes: Determining the loading priorities of the plurality of training data according to the importance scores of the plurality of training data respectively; The target available cache nodes corresponding to the plurality of training data are controlled in sequence, and the plurality of training data are preloaded according to a loading priority order of the plurality of training data.

3. The method according to claim 1, characterized in that The node information includes cached data in the available cache nodes, and determining at least one candidate available cache node from the plurality of available cache nodes according to the training data includes: In a case where the cached data in the available cache node does not match the training data, the available cache node is determined to be a candidate available cache node.

4. The method according to claim 1, wherein The node information includes cached data and remaining cache space in the available cache node, and the method further includes: For any candidate available cache node, when the remaining cache space of the candidate available cache node is less than the data volume of the training data, determining cache data to be removed from the cached data in the candidate available cache node according to the data volume of the training data; The cache data to be removed is removed, and a cache priority of the candidate available cache node when loading the training data is re-determined.

5. The method according to claim 4, characterized in that The cached data includes data corresponding to a plurality of cache tasks, and determining cached data to be removed from the cached data in the candidate available cache nodes according to the data volume of the training data includes: Determining, based on the importance score and data volume corresponding to each cache task, a replacement cost when the data corresponding to each cache task is removed; Determining a removal priority of each cache task according to a replacement cost of each cache task; At least one cache task to be removed is determined from the plurality of cache tasks according to the data volume and removal priority of each cache task, wherein the data corresponding to the at least one cache task to be removed is the cache data to be removed.

6. The method according to claim 5, characterized in that The method further comprises: Determining a replacement cost of removing the at least one cached data to be removed, and determining a corresponding usage value after loading the training data; If the replacement cost of the at least one cache data to be removed is less than the use value, remove the cache data to be removed; or In a case where the replacement cost of the at least one cache data to be removed is greater than or equal to the use value, a target available cache node is re-determined from at least one candidate available cache node according to the cache priority.

7. A training data preloading device, characterized in that: include: An acquisition module, configured to acquire node information of multiple available cache nodes in a target cluster and training data included in the current cache node when a failure of the current cache node is detected; A first determining module is configured to determine cache priorities of the plurality of available cache nodes for loading the training data according to node information of the plurality of available cache nodes respectively; A second determining module is configured to determine at least one candidate available cache node from the plurality of available cache nodes according to the training data; A control module is configured to determine a target available cache node from the at least one candidate available cache node according to the cache priority, and control the target available cache node to preload the training data.

8. A computer device comprising a memory, a processor, and a computer program stored in the memory, wherein: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instruction is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Data set caching method and device, electronic equipment and storage medium

    CN116594923A

  • Cluster-based training method and device, electronic equipment and storage medium

    CN117742959A

  • Task deployment method and device, equipment and storage medium

    CN118295805A

  • Caching method and device, storage medium and electronic equipment

    CN118567791A

  • Storage resource allocation method, storage medium and electronic equipment

    CN120528933A