Parameter determination method and device for model training, equipment and storage medium

By determining the upper limit of the communication network between the training node and the storage node in the model training scenario of the memory separation, the determination of the batch size of the model parameters is optimized, and the problem of unreasonable batch size determination in the prior art is solved and the model training efficiency is improved.

CN120180125APending Publication Date: 2025-06-20CHINA UNITED NETWORK COMM GRP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510265854.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

In the model training scenario of separation of calculations, the determination of the batch size of the model parameters in the prior art is not reasonable enough, which affects the improvement of model training efficiency.

Method used

By determining the first reference upper limit, that is, the upper limit of the number of training samples that the communication network between the training node and the storage node can transmit in each transmission time window, the target batch size of the training node during the model training process is determined, ensuring that the batch size meets the network transmission capability between nodes.

Benefits of technology

It effectively improves the accuracy of model parameter batch size determination, reduces network congestion or training resource limitations caused by unreasonable batch size settings, thereby improving the training efficiency of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180125A_ABST
    Figure CN120180125A_ABST
Patent Text Reader

Abstract

The invention discloses a parameter determination method and device for model training, equipment and a storage medium, and relates to the technical field of machine learning, the method is applied to a training node, the training node is in communication connection with a storage node, the storage node is used for storing a training sample, and the training node is used for obtaining the training sample from the storage node for model training. Comprising the steps that a first reference upper limit is determined, and the first reference upper limit is used for representing the upper limit of the number of training samples which can be transmitted by a communication network between a training node and a storage node in each transmission time window; and based on the first reference upper limit, determining a target batch size of the training node in the model training process, the target batch size being used for representing the number of training samples obtained from the storage node each time. According to the method, transmission limitation of the transmission network bandwidth of the communication network between the training node and the storage node on the training sample is considered, so that the accuracy of determining the size of the model parameter batch can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of machine learning technology, and in particular, to a method, apparatus, device, and storage medium for determining parameters in model training. Background Art

[0002] In the field of machine learning, the performance of artificial intelligence models is closely related to the scale of training data. As the application scenarios of models become increasingly complex, the data processing requirements of models also face new challenges. On the one hand, due to security and privacy protection, it is difficult for enterprises to migrate and utilize their internal data. On the other hand, the explosive growth of data volume requires AI computing centers to be equipped with additional storage resources while ensuring powerful computing power, significantly increasing the construction cost.

[0003] As an emerging solution, the memory-computation separation technology can separate the storage and computation of data. During the model training process, the training node can directly obtain data from the remote storage device, reducing the storage consumption of data and effectively ensuring the security and privacy of data.

[0004] Currently, in the scenario of model training with memory-computation separation, the determination of the batch size of model parameters is not reasonable enough, which may affect the improvement of model training efficiency. Summary of the Invention

[0005] This application provides a method, apparatus, device, and storage medium for determining parameters in model training, which can effectively improve the accuracy of determining the batch size of model parameters.

[0006] In a first aspect, this application provides a method for determining parameters in model training. This method is applied to a training node, and the training node is communicatively connected to a storage node. The storage node is used to store training samples. The training node is used to obtain training samples from the storage node for model training.

[0007] The method includes: determining a first reference upper limit, where the first reference upper limit is used to represent the upper limit of the number of training samples that can be transmitted by the communication network between the training node and the storage node in each transmission time window. Based on the first reference upper limit, determining the target batch size during the model training process of the training node, where the target batch size is used to represent the number of training samples obtained from the storage node each time.

[0008] The method for determining parameters in model training provided by this application first determines the first reference upper limit, which defines the upper limit of the number of training samples that can be transmitted in each transmission time window of the communication network between the training node and the storage node. Then, the first reference upper limit is incorporated into the process of determining the batch size of model parameters, enabling the training node to determine a batch size of model parameters that fits the network transmission capacity between nodes, thereby reducing situations of network congestion or training resource limitations caused by an unreasonable batch size setting. Compared with the prior art, this method takes into account the network transmission bandwidth limitation between nodes, realizes the combination of network transmission and model training, and can effectively improve the training efficiency of the model.

[0009] A possible implementation manner for determining the first reference upper limit includes: determining the first reference upper limit based on the network transmission bandwidth of the communication network, the duration of the transmission time window, and the size of each training sample.

[0010] A possible implementation manner for determining the first reference upper limit based on the network transmission bandwidth of the communication network, the duration of the transmission time window, and the size of each training sample includes: determining the first reference upper limit according to the following formula:

[0011]

[0012] where B1 represents the first reference upper limit, Rnet represents the network transmission bandwidth, Ttrans represents the duration of the transmission time window, and Dsample represents the size of each training sample.

[0013] A possible implementation manner for determining the target batch size of the training node during model training based on the first reference upper limit includes: determining a second reference upper limit and a third reference upper limit. The second reference upper limit is used to represent the upper limit of the number of training samples that the training node can process within a preset one-iteration training duration, and the third reference upper limit is used to represent the upper limit of the number of training samples that the training node can store during one-iteration training. Determine the target batch size based on the minimum value among the first reference upper limit, the second reference upper limit, and the third reference upper limit.

[0014] A possible implementation manner for determining the target batch size based on the minimum value among the first reference upper limit, the second reference upper limit, and the third reference upper limit includes: selecting the power of 2 that is less than the minimum value and closest to the minimum value from multiple powers of 2 as the target batch size.

[0015] A possible implementation manner for determining the second reference upper limit includes: determining the second reference upper limit based on the number of training samples that the training node can process per unit time and the preset one-iteration training duration.

[0016] A possible implementation method for determining a second reference upper limit based on the number of training samples that a training node can process per unit time and a preset duration of one iteration of training includes:

[0017] Determine the second reference upper limit according to the following formula:

[0018] B2 = Ccom · Titer;

[0019] where B2 represents the second reference upper limit, Ccom represents the number of training samples that a training node can process per unit time, and Titer represents the preset duration of one iteration of training.

[0020] A possible implementation method for determining a third reference upper limit includes: determining the third reference upper limit based on the total memory capacity of the training node, the memory size occupied by the training model, and the memory size occupied by each training sample.

[0021] A possible implementation method for determining a third reference upper limit based on the total memory capacity of the training node, the memory size occupied by the training model, and the memory size occupied by each training sample includes:

[0022] Determine the third reference upper limit according to the following formula:

[0023]

[0024] where B3 represents the third reference upper limit, Mtotal represents the total memory capacity of the training node, Mmodel represents the memory size occupied by the training model, and Mbatch represents the memory size occupied by each training sample.

[0025] In a second aspect, the present application provides an apparatus for determining parameters of model training, and the apparatus includes various functional modules for implementing the method in the first aspect above.

[0026] In a third aspect, the present application provides a computer program product, including: computer instructions; when the computer instructions run on an electronic device, the electronic device is enabled to implement the method described in the first aspect above.

[0027] In a fourth aspect, the present application provides an electronic device, and the electronic device includes: a processor and a memory; the memory stores instructions executable by the processor; when the processor is configured to execute the instructions, the electronic device is enabled to implement the method described in the first aspect above.

[0028] In a fifth aspect, the present application provides a readable storage medium, and the readable storage medium includes: software instructions; when the software instructions run in an electronic device, the electronic device is enabled to implement the method described in the first aspect above.

[0029] The beneficial effects of the second to fifth aspects described above can be referred to those described in the first aspect and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0031] Figure 1 Schematic diagram of the scenario of a method for determining parameters of model training provided by an embodiment of the present application;

[0032] Figure 2 Flowchart of a method for determining parameters of model training provided by an embodiment of the present application;

[0033] Figure 3 Flowchart of a method for determining a target batch size based on multiple reference upper limits provided by an embodiment of the present application;

[0034] Figure 4 Schematic diagram of the composition of a device for determining parameters of model training provided by an embodiment of the present application;

[0035] Figure 5 Schematic diagram of the composition of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0036] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0037] It should be noted that in the embodiments of the present application, words such as "exemplarily" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplarily" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly, using words such as "exemplarily" or "for example" is intended to present relevant concepts in a specific manner.

[0038] To facilitate a clear description of the technical solutions of the embodiments of the present application, in the embodiments of the present application, terms such as "first" and "second" are used to distinguish identical or similar items with basically the same functions and roles. Those skilled in the art can understand that the terms "first", "second", etc. do not limit the quantity and execution order.

[0039] As described in the background art, in the model training scenario of separating storage and computing with data remote access, since the training node and the storage node are communicatively connected, the number of training samples (i.e., batch size) that the training node can obtain depends on the transmission capacity of the communication network. For example, if the network bandwidth is small and the batch size of the model is large, it may lead to the model training being in a long data preparation stage, causing the computing resources to be idle and reducing the training efficiency of the model.

[0040] In view of this, how to determine a reasonable batch size of model parameters in the scenario of separating storage and computing has become an urgent problem to be solved currently.

[0041] Based on this, the present application provides a method for determining model training parameters, which can incorporate the network transmission capacity between nodes into the process of determining the batch size of model parameters, and can effectively improve the accuracy of determining the batch size of model parameters.

[0042] To better understand the embodiments of the present application, the following professional terms involved in the following embodiments are described.

[0043] 1. Batch size: Batch size is a key hyperparameter in the training process of deep learning and machine learning models. During model training, since the training dataset is usually very large, it is impossible to input all data into the model for calculation at one time. Therefore, the dataset needs to be divided into several smaller subsets, and these subsets are the so-called "batches". The batch size refers to the number of samples contained in each such batch.

[0044] 2. Separation of storage and computing: An architectural mode in which the data storage and computing functions are separately deployed, and the storage and computing resources are independently managed and expanded.

[0045] 3. Training node: In the structure of separating storage and computing, the device or node that undertakes the core computing task in model training, determines the model training parameters, requests training samples, and conducts model training.

[0046] 4. Storage node: In the architecture of separating storage and computing, the device or node that is specifically responsible for storing data and providing training samples to the training node upon request.

[0047] 5. Remote Direct Memory Access (RDMA) Network: The RDMA network is not an independent network architecture in the traditional sense. Instead, it is a network composed of a high-speed and low-latency data transmission technology system that is built on the underlying network and uses special hardware and protocols to achieve direct memory access between computers. It can be based on Ethernet and InfiniBand technology. Through hardware such as network interface cards that support RDMA functions, data can be transferred directly between the memories of different devices bypassing the operating system kernel without the need for a large amount of CPU participation, thus greatly improving data transmission efficiency and performance. It is mainly applied to scenarios with extremely high requirements for data transmission speed and latency, such as high-performance computing, big data processing, memory-computation separation architectures, and other fields.

[0048] The method for determining parameters in model training provided by this application can be applied to training nodes. The training nodes are communicatively connected to storage nodes, specifically as Figure 1 shown, including: a training node 110, a storage node 120, and a communication network 130 between the training node 110 and the storage node 120.

[0049] The training node 110 is used to determine a first reference upper limit, which is used to represent the upper limit of the number of training samples that can be transmitted by the communication network 130 between the training node 110 and the storage node 120 in each transmission time window.

[0050] The training node 110 is further used to determine a target batch size during the model training process based on the first reference upper limit. The target batch size is used to represent the number of training samples retrieved from the storage node 120 each time.

[0051] The training node 110 is further used to request training samples from the storage node 120 based on the target batch size.

[0052] The storage node 120 is used to send the training samples to the training node 110 in response to the request information of the training node 110.

[0053] The communication network 130 is used to provide a network transmission channel for data transmission between the training node 110 and the storage node 120.

[0054] The training node 110 can be an electronic device with computing and processing functions such as a computer or a server.

[0055] Among them, the server can be a single server, or can also be a server cluster composed of multiple servers. In some embodiments, the server cluster can also be a distributed cluster. Optionally, the server can also be implemented on a cloud platform. For example, the cloud platform can include private cloud, public cloud, hybrid cloud, community cloud, distributed cloud, inter-cloud, and multi-cloud, etc., or any combination thereof. The embodiments of the present application do not limit the specific device form of the training node.

[0056] The storage node 120 can be an electronic device with data storage and management functions such as a hard disk array, a network attached storage (NAS) device, or a storage server.

[0057] The storage node 120 can be an electronic device to independently complete the data storage work, or can also be composed of multiple electronic devices to form a cluster to meet the needs of large-scale data storage and high-concurrency access. In some embodiments, the cluster can also adopt a distributed deployment method to improve the reliability and scalability of storage. The embodiments of the present application do not limit the specific device form and quantity of the storage node 120.

[0058] In some embodiments, the communication network 130 can be an RDMA network implemented based on Ethernet. The RDMA network can reduce network latency, improve bandwidth for data transmission, enable the training node to quickly obtain stored data, and also reduce CPU overhead, making the CPU more focused on model training.

[0059] In other embodiments, the communication network 130 can also be the following several networks: Fibre Channel network, wireless network, or TCP / IP network, etc.

[0060] As Figure 2 shown, a method for determining parameters of model training provided by the present application can be applied to the above-mentioned training node, and the method specifically includes the following steps:

[0061] S101. Determine the first reference upper limit.

[0062] Among them, the first reference upper limit is used to represent the upper limit of the number of training samples that can be transmitted by the communication network between the training node and the storage node in each transmission time window. The transmission time window refers to the time range for transmitting training samples between the training node and the storage node.

[0063] Specifically, in the process of determining the first reference upper limit, it is necessary to obtain the network transmission bandwidth of the communication network between the training node and the storage node, and calculate the first reference upper limit based on the transmission time window duration for the training node to request training samples from the storage node each time (i.e., the duration of the data preparation phase in each iteration training process of the model). Among them, the network transmission bandwidth can be obtained by the training node through network monitoring of the communication network, and the transmission time window duration can be set by the management user as needed.

[0064] It should be noted that in the communication network, the network transmission bandwidth may be affected by network congestion and decrease. When a large amount of data floods into the communication network simultaneously, the processing capacity of network nodes reaches saturation, and data packets queue up waiting for transmission, which will cause network congestion. In this case, the data transmission delay increases, and the packet loss rate rises, which may lead to a decrease in the actual available network transmission bandwidth of the communication network. Therefore, when the training node obtains the network transmission bandwidth, it can calibrate the actual available network transmission bandwidth based on the network congestion status.

[0065] In some embodiments, the training node can evaluate its congestion degree through the communication parameters of the communication network, adjust the network transmission bandwidth according to the congestion degree, obtain the actual network transmission bandwidth that the training node can actually use to transmit training samples, and calculate the first reference upper limit through the actual network transmission bandwidth.

[0066] A possible implementation method is that the training node can determine whether the communication network is in a congested state by monitoring whether the packet delay and packet loss rate in the communication network exceed the preset thresholds. And adjust the network transmission bandwidth according to the excess amounts of the packet delay and packet loss rate to obtain the actual network transmission bandwidth.

[0067] For example, the preset packet delay threshold of a certain training node is 30 milliseconds, the packet loss rate threshold is 0.5%, and the rated network transmission bandwidth is 10 Gbps. Suppose that during the time period of the previous round of model training iteration, the training node monitored that the average packet delay in the communication network reached 50 milliseconds, the excess amount was 20 milliseconds, and the packet loss rate was 1.5%, and the excess amount was 1%. According to the preset adjustment rules, for every 10 milliseconds exceeding the delay threshold, the bandwidth is reduced by 1 Gbps, and for every 0.5% exceeding the packet loss rate threshold, the bandwidth is also reduced by 1 Gbps. Then, due to the packet delay exceeding the standard, the bandwidth needs to be reduced by 20÷10×1 = 2 Gbps, and due to the packet loss rate exceeding the standard, the bandwidth needs to be reduced by 1÷0.5×1 = 2 Gbps. In total, the bandwidth needs to be reduced by 4 Gbps. Therefore, the actual available network transmission bandwidth is 10 - 4 = 6 Gbps.

[0068] S102. Determine the target batch size of the training node during the model training process based on the first reference upper limit.

[0069] Among them, the target batch size is used to represent the number of training samples fetched from the storage node each time.

[0070] Specifically, in the process of determining the target batch size, the training target batch size can be determined not only based on the first reference upper limit, but also based on multiple reference upper limits (such as the first reference upper limit, the second reference upper limit, and the third reference upper limit). For the process of determining the target batch size based on multiple reference upper limits and the introduction of the second reference upper limit and the third reference upper limit, please refer to steps S301 - 302 below and will not be elaborated here.

[0071] It should be noted that during the model training process, the batch size of model parameters can be either fixed or dynamically changing. If the training node adopts a fixed batch size, steps S101 - 102 can be applied to the initial stage of model training to determine a fixed target batch size. If the training node adopts a dynamically changing batch size, steps S101 - S102 can be applied to the data preparation stage of each iteration during model training to determine the batch size of the current round.

[0072] In some embodiments, when the training node adopts a dynamically changing target batch size, hyperparameters such as the learning rate and regularization parameter associated with it can also be adjusted accordingly based on the change of the target batch size.

[0073] A possible implementation is that when the training node adopts a dynamically changing target batch size, when the target batch size increases, to prevent the model update step size from being too large and missing the optimal solution, the learning rate can be appropriately reduced. At the same time, the regularization intensity can be increased accordingly to avoid overfitting of the model. When the target batch size decreases, to avoid slow convergence of the model, the learning rate can be moderately increased to prompt the model to actively update parameters even with fewer samples. At the same time, the regularization intensity can be reduced accordingly to avoid underfitting of the model.

[0074] It should be understood that when determining the target batch size based on the first reference upper limit, the model training can select an appropriate target batch size according to the network condition of the communication network. When the network transmission bandwidth is sufficient and stable, the first reference upper limit may increase accordingly, and then the target batch size may also increase, and the efficiency of training the model will increase appropriately. When the network transmission bandwidth becomes congested and the bandwidth decreases, the first reference upper limit may decrease accordingly, and then the target batch size may also decrease. At this time, the training samples of the training model may decrease to ensure the continuity and stability of training.

[0075] As can be seen from the above steps S101 - S102, by determining the first reference upper limit, the training node can calculate the network transmission capacity (which can also be referred to as network transmission constraint) of the communication network between the training node and the storage node, and can select a relatively accurate target batch size based on the first reference upper limit. Thus, model training and network transmission can be coordinated to reduce the interference caused by communication network fluctuations to the model training process.

[0076] In some embodiments, the first reference upper limit can be determined according to the network transmission parameters between nodes. In this case, step S101 in the method can specifically include:

[0077] S201. Determine the first reference upper limit based on the network transmission bandwidth of the communication network, the duration of the transmission time window, and the size of each training sample.

[0078] Specifically, multiply the network transmission bandwidth of the communication network by the duration of the transmission time window to obtain the total amount of data transmitted, and then divide the total amount of data by the size of a single training sample to obtain the upper limit of the number of training samples that can be transmitted within the duration of the transmission time window, which is used as the first reference upper limit.

[0079] It should be noted that the unit of network transmission bandwidth is generally measured in the number of bits transmitted per second, and its unit can be megabits per second (Mbps), gigabits per second (Gbps), or terabits per second (Tbps), which represents the amount of data that the communication network can transmit per unit time. The duration of the transmission time window is in seconds (s) and is configured by the management user according to the actual requirements of model training and the communication network status. This duration limits the time range of the data preparation phase when the training node requests training samples from the storage node each time. The size of a single training sample is generally in bytes (B). During the calculation process, it is necessary to match the unit of the duration of the transmission time window with the unit of the network transmission bandwidth, and match the size of the training sample with the unit of the network transmission bandwidth.

[0080] In some embodiments, if the training node communicates with multiple storage nodes and the multiple storage nodes are deployed in different geographical locations, the sum of the network transmission bandwidths of the multiple storage nodes can be calculated as the network transmission bandwidth of the communication network between the training node and the multiple storage nodes.

[0081] In some embodiments, the size of the training samples can be standardized before training so that the size of each sample is fixed.

[0082] In some other embodiments, the size of the training samples can also be not fixed. For training samples of different sizes, the approximate fixed-size method can be adopted to determine the size of a single sample.

[0083] For example, there is a set of image training samples in the storage node, with sizes ranging from 100×100 pixels to 800×600 pixels. For training samples of different sizes, the approximate fixed-size method can be adopted to determine the size of a single sample. Through calculation and analysis, all samples are approximated to a size of 1024×1024 pixels to facilitate the subsequent unified training process. After calculation, each such approximately fixed-size image sample occupies about 4MB of memory after specific encoding processing, and this is used as the size of a single sample for subsequent calculations.

[0084] Exemplarily, the calculation process of the first reference upper limit can be expressed as:

[0085]

[0086] where B1 represents the first reference upper limit, Rnet represents the network transmission bandwidth, Ttrans represents the duration of the transmission time window, and Dsample represents the size of each training sample.

[0087] For example, assume that the network transmission bandwidth is 500Mbps, the duration of the transmission time window is 10 seconds, and the size of each training sample is 5MB. First, calculate the product of the network transmission bandwidth and the duration of the transmission time window to obtain the amount of data that can be transmitted within the transmission time window, which is 5000Mb. Convert bits to bytes, which is 5000÷8 = 625MB. Then, divide the above result by the size of a single training sample to calculate the first reference upper limit: 625÷5 = 125.

[0088] In some embodiments, during the process of determining the target batch size in step S102, a second reference upper limit and a third reference upper limit can also be added, as Figure 3 shown, step S102 in this method can specifically include:

[0089] S301. Determine the second reference upper limit and the third reference upper limit.

[0090] The second reference upper limit is used to represent the upper limit of the number of training samples that the training node can process within a preset duration of one iterative training. The third reference upper limit is used to represent the upper limit of the number of training samples that the training node can store during one iterative training process.

[0091] Specifically, the second reference upper limit takes into account the computing power of the training node, and its specific determination method can refer to step S501. The third reference upper limit takes into account the storage capacity of the training node, and its specific determination method can refer to step S601 below. Details are not described here.

[0092] S302. Determine the target batch size based on the minimum value among the first reference upper limit, the second reference upper limit, and the third reference upper limit.

[0093] A possible implementation is to directly use the minimum value among the first reference upper limit, the second reference upper limit, and the third reference upper limit as the target batch size. Ensure that the limits are not exceeded in the three key dimensions of network transmission, computing power, and storage capacity, and simply and directly ensure the smooth progress of the training process. It is applicable to scenarios with extremely high requirements for the stability of model training, relatively balanced resource dimension limits, and no special optimization requirements.

[0094] Exemplarily, the calculation process of step S302 can be expressed as:

[0095] Bmin = min(B1, B2, B3) Formula (2)

[0096] Where B1 represents the first reference upper limit, B2 represents the second reference upper limit, B3 represents the third reference upper limit, and Bmin represents the minimum value among the multiple reference upper limits.

[0097] For example, directly select the minimum value among multiple reference upper limits. Assume that the first reference upper limit is 500 training samples, the second reference upper limit is 450 training samples, and the third reference upper limit is 480 training samples. Then, based on the minimum value of 450 among the multiple reference upper limit values, determine the target batch size.

[0098] Another possible implementation is that before determining the target batch size, a fault tolerance analysis can be performed on each reference upper limit to estimate possible resource fluctuations or short-term failures and adjust the reference upper limits. For network resources, consider that network congestion may occur in the communication network, and reserve a certain amount of redundant transmission bandwidth when calculating based on the first reference upper limit. For computing resources, consider situations such as temporary hardware frequency reduction and appropriately reduce the second reference upper limit. For storage resources, consider the stability of data storage and set a certain safety margin for the third reference upper limit. Finally, select the minimum value among these adjusted reference upper limits as the target batch size to enhance the adaptability of the training process to unexpected situations and ensure the stability and reliability of the training.

[0099] For example, after performing fault tolerance analysis and adjustment on the reference upper limit, the minimum value among them is selected. Suppose the original first reference upper limit is 600 training samples, which is calculated under the condition of relatively ideal network transmission bandwidth. However, considering that there may be momentary congestion in network transmission, such as during peak office hours, a large amount of data transmission by other devices in the network may preempt the bandwidth. Therefore, a 30% redundancy space is reserved for the first reference upper limit, and the adjusted first reference upper limit becomes 600×(1 - 30%) = 420. The original second reference upper limit is 550 training samples. However, the hardware of the training node may experience temporary frequency reduction during long-term high-load operation. Suppose it is estimated that the computing power will be reduced by 15%. Then the adjusted second reference upper limit is 550×(1 - 15%) = 467.5, which is rounded down to 467. The original third reference upper limit is 520 training samples. To ensure the stability of data storage and avoid read and write errors due to storage approaching full load, a 20% safety margin is set, and the adjusted third reference upper limit is 520×(1 - 20%) = 416. After fault tolerance analysis and adjustment, comparing the three adjusted reference upper limits, the minimum value can be obtained as 450.

[0100] It should be understood that through fault tolerance analysis of the reference upper limit, the reference upper limit can be made more in line with the actual scenario, so that the determination of the target batch size is more reliable and stable.

[0101] Another possible implementation method is based on the operation principle of the graphics processing unit (GPU) in the training node. When the batch size is a power of 2, data can be more evenly distributed to each processing core. In this case, step S302 above can specifically include:

[0102] S401. Select the power of 2 that is less than the minimum value and closest to the minimum value from multiple powers of 2 as the target batch size.

[0103] In some embodiments, a set of values of powers of 2 can be preset during the training phase. After determining the minimum value among multiple reference upper limits, the binary search method is used to find the power of 2 that is less than the minimum value and closest to it in the value set as the target batch size.

[0104] For example, for the selected minimum value, find the power of 2 closest to it downward as the target batch size. Suppose the minimum value selected from multiple reference upper limits is 150, and the preset set of power-of-2 values is [1, 2, 4, 8, 16, 32, 64, 128, 256, 512, 1024]. Using the binary search method, first locate the head and tail of the set, take the middle value 32 and compare it with the minimum value 150. Since 32 is less than 150, take the subset 1 [32, 64, 128, 256, 512, 1024] of the numerical set as the new search range. Then locate the head and tail of subset 1, take the middle value 256 and compare it with the minimum value 150. Since 256 is greater than 150, take the subset 2 [32, 64, 128] of subset 1 as the new range. Then locate the head and tail of subset 2, take the middle value 64 and compare it with the minimum value 150. Since 64 is less than 150, take the subset 3

[128] of subset 2 as the new range. At this time, there is only one value 128 in this subset. Since 128 is less than 150 and is the closest value less than 150 that can be found currently, 128 is used as the target batch size.

[0105] It can be understood that step S401 utilizes the computing characteristics of the processor to calculate the target batch size equal to the power of 2, enabling the training sample data to be evenly distributed to each processing core, making full use of GPU resources, and thus improving the computing parallelism and model training efficiency.

[0106] As can be seen from the above steps S301 - S302, the first reference upper limit considers the limitation of the network resources between the training node and the storage node on the batch size, the second reference upper limit considers the limitation of the computing resources of the training node on the batch size, and the third reference upper limit considers the limitation of the storage resources of the training node on the batch size. By comprehensively considering these three different levels of reference upper limits and selecting the minimum value among them as the target batch size, it is possible to effectively balance the resource consumption in network transmission, computing, and storage, avoid training bottlenecks caused by insufficient resources in a certain aspect, and ensure the smoothness and stability of the entire training process.

[0107] The following is an introduction to the determination process of the second reference upper and lower limits and the third reference upper limit in the embodiments of the present application.

[0108] In some embodiments, the second reference upper limit in step S301 can be determined based on the computing power of the training node. Step S301 can specifically include:

[0109] S501. Determine the second reference upper limit based on the number of training samples that the training node can process per unit time and the preset one - iteration training duration.

[0110] Specifically, as can be seen from the foregoing, the second reference upper limit is used to represent the upper limit of the number of training samples that a training node can process within a preset one-iteration training duration. That is, the second reference upper limit takes into account the computing power of the training node.

[0111] In a possible implementation, the number of sample trainings that a training node can process per unit time can be estimated based on the number of cores of the GPU, the video memory bandwidth, the clock frequency, as well as the number of cores, the main frequency, and the cache size of the central processing unit (CPU). Substituting these performance parameters into a specific calculation model or empirical formula can estimate the number of sample trainings that the training node can process per unit time. Then, combined with the preset one-iteration training duration, multiplying the two can obtain the upper limit of the number of training samples that the training node can process within the preset one-iteration training duration, that is, determine the second reference upper limit.

[0112] The process of the second reference upper limit can be expressed as:

[0113] B2 = Ccom · Titer Formula (3)

[0114] Wherein, B2 represents the second reference upper limit, Ccom represents the number of training samples that the training node can process per unit time, and Titer represents the preset one-iteration training duration.

[0115] For example, assume that the number of cores of the GPU of a certain training node is 8000, the video memory bandwidth is 600 GB / s, the clock frequency is 1.8 GHz, the number of cores of the CPU is 32, the main frequency is 3.5 GHz, and the cache size is 64 MB. Substitute into the empirical formula designed specifically for this type of hardware and training task: the number of samples processed per unit time = (number of GPU cores × 0.01 + video memory bandwidth × 0.005 + clock frequency × 10 + number of CPU cores × 5 + main frequency × 2 + cache size × 0.1). After calculation, the number of samples processed per unit time (per second) is (8000 × 0.01 + 600 × 0.005 + 1.8 × 10 + 32 × 5 + 3.5 × 2 + 64 × 0.1) = 275.4, approximately 275. If the preset one-iteration training duration is 100 seconds, then the second reference upper limit is 275 × 100 = 27500 training samples. It should be noted that the above empirical formula can be used as an example of a reference method for obtaining the number of training samples that a training node can process per unit time. The embodiments of the present application do not limit the specific method for obtaining the number of training samples that a training node can process per unit time.

[0116] In some embodiments, the third reference upper limit in step S301 can be determined based on the storage capacity of the training node. Step S301 can also specifically include:

[0117] S601. Determine a third reference upper limit based on the total memory capacity of the training node, the memory size occupied by the training model, and the memory size occupied by each training sample.

[0118] Specifically, as can be seen from the foregoing, the third reference upper limit is used to represent the upper limit of the number of training samples that can be stored by the training node during one iteration of training. That is, the second reference upper limit takes into account the memory capacity of the training node.

[0119] Specifically, the total memory capacity of the training node can be obtained by calling a tool or command to obtain the memory information of the system. The memory size occupied by the training model can be calculated by recording the memory size of the system before the training starts and the memory size of the system after the training starts, and calculating the difference between the two as the memory size occupied by the model.

[0120] In some embodiments, the operating system of the training node itself needs to occupy a certain amount of memory to run, that is, the memory size occupied by the necessary programs of the system. In actual calculations, the memory size occupied by the necessary programs can be regarded as a part of the space occupied during the operation of the model, because in the method of obtaining the memory size occupied by the model, the memory size occupied by the operating system may be calculated together.

[0121] The process of the third reference upper limit can be expressed as:

[0122]

[0123] Among them, B3 represents the third reference upper limit, Mtotal represents the total memory capacity of the training node, Mmodel represents the memory size occupied by the training model, and Mbatch represents the memory size occupied by each training sample.

[0124] For example, assume that the training node obtains the total memory capacity of the training node as 32 GB through the system command. Before the training starts, record the system memory usage as 2 GB (that is, the memory size occupied by the operating system). After the training model is started, loaded, and run for a period of time, record the system memory usage as 6 GB. Then the memory size occupied by the training model is 6 GB (the sum of the memory size occupied by the operating system and the memory size occupied by the model). Assume that the training sample is processed image data, and the memory size occupied by each training sample is fixed at 0.01 GB (that is, 10.24 MB). Then substitute the values into formula (4) to obtain the third reference upper limit 32 - 6 / 0.01 = 2600.

[0125] The above mainly introduced the solution provided by the embodiments of the present application from the perspective of methods. To implement the above functions, it includes the corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should easily realize that, combining the units and algorithm steps of each example described in the embodiments disclosed in this article, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.

[0126] In an exemplary embodiment, the embodiments of the present application further provide a device for determining parameters for model training, and this device can be applied to the above training node. As Figure 4 shown, this device includes: a determination module 410 and a processing module 420.

[0127] The determination module 410 is used to determine a first reference upper limit. The first reference upper limit is used to represent the upper limit of the number of training samples that can be transmitted by the communication network between the training node and the storage node in each transmission time window.

[0128] The processing module 420 is used to determine the target batch size of the training node during the model training process based on the first reference upper limit. The target batch size is used to represent the number of training samples obtained from the storage node each time.

[0129] A possible implementation manner is that the determination module 410 is specifically used to determine the first reference upper limit based on the network transmission bandwidth of the communication network, the duration of the transmission time window, and the size of each training sample.

[0130] A possible implementation manner is that the determination module 410 is specifically used to determine the first reference upper limit according to the following formula:

[0131]

[0132] Among them, B1 represents the first reference upper limit, Rnet represents the network transmission bandwidth, Ttrans represents the duration of the transmission time window, and Dsample represents the size of each training sample.

[0133] A possible implementation manner is that the processing module 420 is specifically used to determine a second reference upper limit and a third reference upper limit. The second reference upper limit is used to represent the upper limit of the number of training samples that the training node can process within a preset one-iteration training duration, and the third reference upper limit is used to represent the upper limit of the number of training samples that the training node can store during one-iteration training. Based on the minimum value among the first reference upper limit, the second reference upper limit, and the third reference upper limit, determine the target batch size.

[0134] In a possible implementation, the processing module 420 is specifically configured to select, from multiple powers of 2, the power of 2 that is less than and closest to the minimum value as the target batch size.

[0135] In a possible implementation, the determining module 410 is further configured to determine a second reference upper limit based on the number of training samples that a training node can process per unit time and a preset duration of one iteration of training.

[0136] In a possible implementation, the determining module 410 is further configured to determine the second reference upper limit according to the following formula:

[0137] B2 = Ccom · Titer;

[0138] where B2 represents the second reference upper limit, Ccom represents the number of training samples that a training node can process per unit time, and Titer represents the preset duration of one iteration of training.

[0139] In a possible implementation, the determining module 410 is further configured to determine a third reference upper limit based on the total memory capacity of the training node, the memory size occupied by the training model, and the memory size occupied by each training sample.

[0140] In a possible implementation, the determining module 410 is further configured to determine the third reference upper limit according to the following formula:

[0141]

[0142] where B3 represents the third reference upper limit, Mtotal represents the total memory capacity of the training node, Mmodel represents the memory size occupied by the training model, and Mbatch represents the memory size occupied by each training sample.

[0143] It should be noted that Figure 4 the division of modules herein is illustrative, merely a logical function division, and there may be other division methods in actual implementation. For example, two or more functions can also be integrated into one processing module. The above integrated module can be implemented in the form of hardware or in the form of a software function module.

[0144] In an exemplary embodiment, as described above, the computing device may specifically be an electronic device with computing and processing capabilities such as a computer or a service. In this case, the embodiments of the present application further provide an electronic device, Figure 5 which is a schematic diagram of the composition of an electronic device provided by the embodiments of the present application. As Figure 5As shown, the electronic device includes: a processor 10, a memory 20, a communication line 30, a communication interface 40, and an input / output interface 50.

[0145] Among them, the processor 10, the memory 20, the communication interface 40, and the input / output interface 50 can be connected through the communication line 30.

[0146] The processor 10 is configured to execute the instructions stored in the memory 20 to implement the method for determining the parameters of model training provided in the foregoing embodiments of the present application. The processor 10 can be a CPU, a general-purpose processor, a network processor (NP), a digital signal processor (DSP), a microprocessor, a micro control unit (MCU) / single-chip microcomputer / single-chip microcontroller, a programmable logic device (PLD), or any combination thereof. The processor 10 can also be any other device with processing functions, such as a circuit, a device, or a software module, which is not limited in the embodiments of the present application. In one example, the processor 10 can include one or more CPUs, such as Figure 5 CPU0 and CPU1 in. As an optional implementation, the electronic device can include multiple processors. For example, in addition to the processor 10, it can also include a processor 60 ( Figure 5 illustrated by a dashed line in).

[0147] The memory 20 is used to store instructions. For example, the instructions can be a computer program. Optionally, the memory 20 can be a read-only memory (ROM) or other types of static storage devices that can store static information and / or instructions, or a random access memory (RAM) or other types of dynamic storage devices that can store information and / or instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, which is not limited in the embodiments of the present application.

[0148] It should be noted that the memory 20 can exist independently of the processor 10 or be integrated with the processor 10. The memory 20 can be located inside the electronic device or outside the electronic device, and the embodiments of the present application do not limit this.

[0149] The communication line 30 is used to transmit information between the components included in the electronic device.

[0150] The communication interface 40 is used to communicate with other devices or other communication networks. The other communication network can be an Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc. The communication interface 40 can be a module, a circuit, a transceiver, or any device capable of implementing communication.

[0151] The input / output interface 50 is used to implement the human-computer interaction between the user and the electronic device. For example, it realizes the action interaction or information interaction between the user and the electronic device.

[0152] Exemplarily, the input / output interface 50 can be a mouse, a keyboard, a display screen, or a touch display screen, etc. Through a mouse, a keyboard, a display screen, or a touch display screen, etc., the action interaction or information interaction between the user and the electronic device can be realized.

[0153] It should be noted that Figure 5 the structure shown in Figure 5 does not constitute a limitation on the electronic device. In addition to

[0154] the components shown, the electronic device can include more or fewer components than those shown in the figure, or a combination of certain components, or a different component arrangement.

[0155] In an exemplary embodiment, the embodiments of the present application further provide a computer program product, which includes computer instructions. When the computer instructions run in the electronic device, the electronic device implements the method in the foregoing method embodiments.

[0156] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using a software program, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer-executable instructions. When the computer-executable instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer-executable instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer-executable instructions can be transmitted from a website, a computer, a server, or a data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.).

[0157] Although the present application has been described in conjunction with various embodiments, however, in the process of implementing the claimed present application, those skilled in the art can understand and implement other variations of the disclosed embodiments by viewing the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "one" does not exclude a plurality. A single processor or other unit can implement several functions recited in the claims. Certain measures are recited in mutually different dependent claims, but this does not mean that these measures cannot be combined to produce good results.

[0158] Although the present application has been described in conjunction with specific features and their embodiments, it is obvious that various modifications and combinations can be made without departing from the spirit and scope of the present application. Accordingly, the present specification and the drawings are only exemplary descriptions of the present application defined by the appended claims, and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of the present application. Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these changes and modifications.

[0159] As described above, the above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for determining parameters for model training, characterized in that: The method is applied to a training node; The training node is communicatively connected with the storage node; The storage node is used to store training samples; The training node is used to obtain training samples from the storage node for model training; the method includes: Determine a first reference upper limit; the first reference upper limit is used to represent the upper limit of the number of training samples that can be transmitted by the communication network between the training node and the storage node in each transmission time window; Based on the first reference upper limit, a target batch size of the training node during the model training process is determined; the target batch size is used to represent the number of training samples obtained from the storage node each time.

2. The method according to claim 1, characterized in that The determining of the first reference upper limit comprises: The first reference upper limit is determined based on a network transmission bandwidth of the communication network, a duration of the transmission time window, and a size of each training sample.

3. The method according to claim 2, characterized in that The determining the first reference upper limit based on the network transmission bandwidth of the communication network, the duration of the transmission time window, and the size of each training sample includes: The first reference upper limit is determined according to the following formula: Among them, B1 represents the first reference upper limit; Rnet represents the network transmission bandwidth; Ttrans represents the duration of the transmission time window; and Dsample represents the size of each training sample.

4. The method according to claim 1, characterized in that The determining, based on the first reference upper limit, a target batch size of the training node in the model training process includes: Determine a second reference upper limit and a third reference upper limit; the second reference upper limit is used to indicate the upper limit of the number of training samples that the training node can process within a preset one-iteration training duration; the third reference upper limit is used to indicate the upper limit of the number of training samples that the training node can store during one-iteration training process; The target batch size is determined based on a minimum value among the first reference upper limit, the second reference upper limit, and the third reference upper limit.

5. The method according to claim 4, characterized in that The determining the target batch size based on a minimum value among the first reference upper limit, the second reference upper limit, and the third reference upper limit comprises: A power of 2 that is smaller than the minimum value and closest to the minimum value is selected from multiple powers of 2 as the target batch size.

6. The method according to claim 4, characterized in that The determining of the second reference upper limit comprises: The second reference upper limit is determined based on the number of training samples that can be processed by the training node within a unit time and the preset training duration for one iteration.

7. The method according to claim 6, characterized in that The determining the second reference upper limit based on the number of training samples that can be processed by the training node within a unit time and the preset one-iteration training duration includes: The second reference upper limit is determined according to the following formula: B2=Ccom·Titer; Among them, B2 represents the second reference upper limit; Ccom represents the number of training samples that the training node can process per unit time; Titer represents the preset duration of one iteration training.

8. The method according to claim 4, characterized in that The determining of the third reference upper limit comprises: The third reference upper limit is determined based on the total memory capacity of the training node, the memory size occupied by the training model, and the memory size occupied by each training sample.

9. The method according to claim 8, characterized in that The determining the third reference upper limit based on the total memory capacity of the training node, the memory size occupied by the training model, and the memory size occupied by each training sample includes: The third reference upper limit is determined according to the following formula: Among them, B3 represents the third reference upper limit; Mtotal represents the total memory capacity of the training node; Mmodel represents the memory size occupied by the training model; Mbatch represents the memory size occupied by each training sample.

10. A parameter determination device for model training, characterized in that: The device comprises: a determination module and a processing module; A determination module, configured to determine a first reference upper limit, where the first reference upper limit is used to represent an upper limit on the number of training samples that can be transmitted by the communication network between the training node and the storage node in each transmission time window; A processing module is used to determine a target batch size of the training node during the model training process based on the first reference upper limit; the target batch size is used to represent the number of training samples obtained from the storage node each time.

11. An electronic device, characterized in that: include: Processor and memory; The memory stores instructions executable by the processor; When the processor is configured to execute the instructions, the electronic device implements the method according to any one of claims 1 to 9.

12. A readable storage medium, characterized in that: include: Software instructions; When the software instructions are executed in an electronic device, the electronic device implements the method according to any one of claims 1 to 9.

13. A computer program product, characterized in that include: Computer instructions; When the computer instructions are executed in an electronic device, the electronic device is enabled to implement the method according to any one of claims 1 to 9.