Calculation parameter acquisition method and device based on DeepSpeed framework and storage medium
By optimizing the storage and loading methods of parameters in the DeepSpeed framework, the problem of low parameter calculation efficiency during model training is solved, and the graphics card is more efficiently calculated and calculated waiting time is reduced.
Patent Information
- Application Number
- CN202510276480.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-13
AI Technical Summary
When training the model under the DeepSpeed framework, the calculation efficiency of the parameters is low, mainly because a large number of model parameters occupy the video memory capacity of the image processing unit (GPU), resulting in a decrease in computing efficiency.
A calculation parameter acquisition method based on the DeepSpeed framework is proposed. By determining the target parameter file used for forward propagation of the current calculation time, and determining whether it is stored in locked memory or cache area. If stored in the cache area, it is sent to the video memory and the parameter file is read from the expansion card to the cache area to avoid excessive use of the video memory.
By optimizing the storage and loading methods of parameters, excessive use of video memory is avoided, allowing the graphics card to perform calculations more efficiently, and at the same time, reading the possible parameter files from the expansion card to the cache area in advance, reducing the calculation waiting time.
Smart Images

Figure CN120144303A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and particularly to a method, device, and storage medium for obtaining calculation parameters based on the DeepSpeed framework. Background Art
[0002] As a framework for optimizing the training of ultra-large-scale models, DeepSpeed requires a large number of model parameters during the forward propagation calculation process. However, the video memory capacity of a graphics processing unit (GPU) is relatively limited, and the storage of a large number of model parameters in the video memory of the graphics processing unit will occupy the video memory capacity, thereby resulting in low calculation efficiency of the model parameters.
[0003] The above content is only used to assist in understanding the technical solution of this application, and does not represent an admission that the above content is prior art. Summary of the Invention
[0004] The main purpose of this application is to provide a method, device, and storage medium for obtaining calculation parameters based on the DeepSpeed framework, aiming to solve the technical problem of low calculation efficiency of parameters during model training under the DeepSpeed framework.
[0005] To achieve the above purpose, this application proposes a method for obtaining calculation parameters based on the DeepSpeed framework, and the method includes: Determine the target parameter file for forward propagation calculation at the current calculation moment; Judge whether the target parameter file is stored in the locked memory or the buffer; If the target parameter file is stored in the buffer, send the target parameter file in the buffer to the video memory, and read the parameter file from the expansion card to the buffer.
[0006] In an embodiment, before the step of determining the target parameter file for forward propagation calculation at the current calculation moment, the method further includes: Based on the capacity of the locked memory and the capacity of all parameter files corresponding to the forward propagation calculation, determine the first parameter file to be unloaded to the locked memory and the number of the first parameter files; Based on the number of the first parameter files, unload the first parameter files from the expansion card to the locked memory.
[0007] In an embodiment, after the step of judging whether the target parameter file is stored in the locked memory or the buffer, the method further includes: If the target parameter file is stored in the locked memory, read the target parameter file in the locked memory into the video memory, and read the parameter file from the expansion card into the buffer.
[0008] In one embodiment, after the step of determining whether the target parameter file is stored in the locked memory or the buffer, the method further includes: If the target parameter file is not stored in the locked memory or the buffer, after reading the target parameter file from the expansion card into the buffer, execute the step of sending the target parameter file in the buffer to the video memory, and reading the parameter file from the expansion card into the buffer.
[0009] In one embodiment, the step of reading the parameter file from the expansion card into the buffer includes: Determine the parameter file to be read in the expansion card based on the calculation execution order of the parameter file; If the parameter file to be read is not stored in the locked memory or the buffer, read the parameter file to be read from the expansion card into the buffer.
[0010] In one embodiment, the step of reading the parameter file to be read from the expansion card into the buffer includes: Determine the number of parameter files to be read based on the remaining storage capacity of the buffer and the capacity of the parameter file to be read; Read the parameter file to be read into the buffer based on the number of parameter files to be read and the calculation execution order of the parameter file.
[0011] In one embodiment, the method further includes: After the forward propagation calculation at the current calculation moment is completed, determine whether the target parameter file is stored in the buffer; If the target parameter file is stored in the buffer, delete the target parameter file in the buffer.
[0012] In one embodiment, the step of determining the first parameter file to be unloaded to the locked memory and the number of the first parameter files based on the capacity of the locked memory and the capacity of all the parameter files corresponding to the forward propagation calculation includes: Obtain the historical calculation frequency of all the parameter files corresponding to the forward propagation calculation; Determine the unloading weight value of the parameter file based on the historical calculation frequency and capacity of the parameter file; Determine the first parameter file and the number of the first parameter files among all parameter files based on the offloading weight value of the parameter file, the capacity of the locked memory, and the capacity of the parameter file.
[0013] In addition, to achieve the above object, the present application also provides a computing parameter acquisition device based on the DeepSpeed framework. The device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor. The computer program is configured to implement the steps of the computing parameter acquisition method based on the DeepSpeed framework as described above.
[0014] In addition, to achieve the above object, the present application also provides a storage medium. The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the computing parameter acquisition method based on the DeepSpeed framework as described above.
[0015] The present application provides a computing parameter acquisition method based on the DeepSpeed framework. First, determine the target parameter file for forward propagation calculation at the current calculation moment, and determine whether the target parameter file for computing parameter acquisition of the framework is stored in the locked memory or the buffer area. If the target parameter file for computing parameter acquisition of the framework is stored in the buffer area for computing parameter acquisition of the framework, then send the target parameter file for computing parameter acquisition in the buffer area for computing parameter acquisition of the framework to the video memory, and read the parameter file from the expansion card to the buffer area for computing parameter acquisition of the framework.
[0016] By storing the parameter file in the locked memory, buffer area, and expansion card, when the target parameter file is needed at the current calculation moment, after determining whether the target parameter file is in the locked memory or the buffer area, the parameter is loaded from the locked memory or the buffer area to the video memory, avoiding excessive occupation of the video memory and enabling the graphics card to perform calculations more efficiently. At the same time, read the new parameter file from the expansion card to the buffer area to prepare in advance for the parameter files that may be needed later, effectively reducing the calculation waiting time. Description of the Drawings
[0017] The drawings here are incorporated into the specification and form a part of the specification, showing the embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0018] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0019] Figure 1 This is a schematic flowchart provided for the first embodiment of the method for obtaining calculation parameters based on the DeepSpeed framework in this application; Figure 2 This is a schematic diagram for obtaining calculation parameters provided for the first embodiment of the method for obtaining calculation parameters based on the DeepSpeed framework in this application; Figure 3 This is a schematic flowchart provided for the second embodiment of the method for obtaining calculation parameters based on the DeepSpeed framework in this application; Figure 4 This is a schematic flowchart of the brief method involved in the method for obtaining calculation parameters based on the DeepSpeed framework in the embodiments of this application; Figure 5 This is a schematic diagram of the device structure of the hardware operating environment involved in the method for obtaining calculation parameters based on the DeepSpeed framework in the embodiments of this application.
[0020] The realization of the purpose, functional features, and advantages of this application will be further described in combination with the embodiments with reference to the accompanying drawings. Specific Embodiments
[0021] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of this application and are not used to limit this application.
[0022] To better understand the technical solutions of this application, the following will be described in detail in combination with the accompanying drawings of the specification and specific embodiments.
[0023] As a framework for optimizing the training of ultra-large-scale models, DeepSpeed requires a large number of model parameters during the forward propagation calculation process. However, the video memory capacity is relatively limited, and the storage of a large number of model parameters in the video memory will occupy the video memory capacity, thereby resulting in low calculation efficiency of the model parameters.
[0024] In view of the above problems, this application proposes a method for obtaining calculation parameters based on the DeepSpeed framework. First, determine the target parameter file for forward propagation calculation at the current calculation moment, and judge whether the target parameter file for obtaining calculation parameters of the framework is stored in the locked memory or cache area. If the target parameter file for obtaining calculation parameters of the framework is stored in the cache area for obtaining calculation parameters of the framework, then send the target parameter file for obtaining calculation parameters of the framework in the cache area for obtaining calculation parameters of the framework to the video memory, and read the parameter file from the expansion card to the cache area for obtaining calculation parameters of the framework.
[0025] The above method stores the parameter file in the locked memory, buffer, and expansion card, and only when the target parameter file is needed at the current calculation moment, after determining whether the target parameter file is in the locked memory or buffer, the parameter is loaded from the locked memory or buffer to the video memory, avoiding excessive occupation of the video memory and enabling the graphics card to perform calculations more efficiently. At the same time, a new parameter file is read from the expansion card to the buffer to prepare in advance for the parameter file that may be needed later, effectively reducing the calculation waiting time.
[0026] It should be noted that the execution entity of this embodiment can be a computing service device with data processing, network communication, and program running functions, such as a tablet computer, a personal computer, etc., or an electronic device capable of implementing the above functions. To more clearly elaborate this technical solution, this application is described with a parameter reading and computing system as the execution entity.
[0027] Based on this, the first embodiment proposed in this application provides a method for obtaining computing parameters based on the DeepSpeed framework. Refer to Figure 1 , in this embodiment, the method for obtaining computing parameters based on the DeepSpeed framework includes steps S10 to S30: Step S10, determine the target parameter file for forward propagation calculation at the current calculation moment.
[0028] Step S20, determine whether the target parameter file is stored in the locked memory or buffer.
[0029] Step S30, if the target parameter file is stored in the buffer, send the target parameter file in the buffer to the video memory, and read a parameter file from the expansion card to the buffer.
[0030] It should be noted that in deep learning, a parameter file refers to the trainable variables in a model, such as weights and biases in a neural network. These parameter files are continuously updated during the training process to minimize the loss function of the model. Among them, the locked memory is pre-allocated and does not exchange data with other storage areas such as the video memory and expansion card. The buffer is a memory area for temporarily storing data and serves as a bridge for the transfer of parameter files between the expansion card and the video memory. Exemplarily, in the specific application scenario of forward propagation calculation, the locked memory and buffer can be different storage areas in the CPU memory, and the expansion card can be a disk.
[0031] At the nth computing moment, the system determines the target parameter file required at the current moment according to the requirements of the current forward propagation calculation. Then, it monitors the storage status of the locked memory and the buffer in real time. Specifically, it determines whether the target parameter file is fixedly stored in the locked memory or has been previously read from the expansion card to the buffer. When the target parameter file is in the buffer, the system sends the target parameter file to the video memory through the data transmission interface and uses the target parameter file in the video memory to perform the forward propagation calculation of the current step. While sending the target parameter file in the buffer to the video memory, the system reads a new parameter file from the expansion card to the buffer to pre-load the parameter files that may be required subsequently into the buffer, reducing the waiting time.
[0032] Optionally, if the target parameter file is stored in the locked memory, the target parameter file in the locked memory is read into the video memory, and a parameter file is read from the expansion card to the buffer; if the target parameter file is not stored in the locked memory or the buffer, after reading the target parameter file from the expansion card to the buffer, the steps of sending the target parameter file in the buffer to the video memory and reading a parameter file from the expansion card to the buffer are performed.
[0033] Exemplarily, referring to Figure 2 , if the target parameter file for forward propagation calculation at the current computing moment is parameter file 1, after determining that parameter file 1 is stored in the locked memory, parameter file 1 is sent from the locked memory to the video memory. At the same time, other parameter files in the expansion card are read into the buffer. If the target parameter file for forward propagation calculation at the current computing moment is parameter file 2, after determining that parameter file 2 is not stored in the locked memory or the buffer, parameter file 2 is first read from the expansion card to the buffer, and then the parameter file 2 in the buffer is sent to the video memory. At the same time, other parameter files in the expansion card are read into the buffer. If the target parameter file for forward propagation calculation at the current computing moment is parameter file 3, after determining that parameter file 3 is stored in the buffer, the parameter file 3 in the buffer is sent to the video memory. At the same time, other parameter files in the expansion card are read into the buffer.
[0034] It can be understood that in the training of traditional deep learning models, forward propagation calculation is one of the core steps. The parameter files required for model training are usually stored in the video memory so that the graphics card can quickly access and process these data. However, the capacity of the video memory is limited. When there are too many parameter files, the video memory may be filled up, causing the graphics card to be unable to perform calculations efficiently. In this case, the forward propagation calculation speed will decrease significantly, and even the calculation process may not be able to continue. In this embodiment, the parameter files are stored in the locked memory, buffer, and expansion card. When the target parameter file is needed at the current calculation moment, after determining whether the target parameter file is in the locked memory or buffer, the parameter is loaded from the locked memory or buffer to the video memory, avoiding over-occupation of the video memory and enabling the graphics card to perform calculations more efficiently. At the same time, the system will also read new parameter files from the expansion card to the buffer to prepare for the parameter files that may be needed subsequently, effectively reducing the calculation waiting time.
[0035] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar content as that in the above-mentioned first embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 3 , before step S10, the method for obtaining calculation parameters based on the DeepSpeed framework further includes steps S40 to S50: Step S40, based on the capacity of the locked memory and the capacities of all parameter files corresponding to the forward propagation calculation, determine the first parameter file to be unloaded to the locked memory and the number of the first parameter files.
[0036] Step S50, based on the number of the first parameter files, unload the first parameter files from the expansion card to the locked memory.
[0037] Optionally, obtain the historical calculation frequencies of all parameter files corresponding to the forward propagation calculation, determine the unloading weight values of the parameter files based on the historical calculation frequencies and capacities of the parameter files, and determine the first parameter file and the number of the first parameter files among all the parameter files based on the unloading weight values of the parameter files, the capacity of the locked memory, and the capacities of the parameter files.
[0038] Exemplarily, in order to store more frequently computed and larger-capacity parameter files in the locked memory in advance, an offloading weight value for each parameter file is calculated based on the historical computation frequency and capacity of the parameter file. The offloading weight value can be defined as the weighted sum of the historical computation frequency and capacity. For example, offloading weight value = α × historical computation frequency + β × capacity, where α and β are weight coefficients. All parameter files are sorted according to the calculated offloading weight values. Among them, the parameter file with a higher offloading weight value is preferentially offloaded. From the sorted parameter file list, the first parameter file is selected in descending order of the offloading weight value and offloaded to the locked memory until the remaining capacity of the locked memory is less than the next parameter file.
[0039] It can be understood that preferentially offloading parameter files with a higher computation frequency and larger capacity to the locked memory can avoid repeatedly reading these larger-capacity parameter files from the expansion card to the buffer area during subsequent forward propagation operations, reduce the data transfer time, and improve the forward propagation calculation efficiency.
[0040] Optionally, the capacity of each parameter file is obtained, and the first parameter file is randomly selected from the parameter files and offloaded from the expansion card to the locked memory, where the total capacity of the first parameter file is less than or equal to the capacity of the locked memory.
[0041] Exemplarily, an empty selected file list and a cumulative capacity variable are initialized. A parameter file is randomly selected, added to the selected file list, and the cumulative capacity is increased by the capacity of the selected parameter file. The above selection steps are repeated until the cumulative capacity reaches a preset threshold. For example, the preset threshold can be a preset ratio of the capacity of the locked memory.
[0042] Optionally, based on the calculation execution order of the parameter files, the first parameter file is determined from the parameter files at a preset interval, and the number of the first parameter files is determined based on the capacity of the parameter files and the capacity of the locked memory.
[0043] Exemplarily, the parameter files are arranged in the calculation execution order to form a parameter file list. In this parameter file list, starting from the first parameter file, the list is traversed in the calculation execution order, and parameter files are selected according to the preset interval to form a new parameter file list. At the same time, based on the capacity of the parameter files in the new parameter file list and the capacity of the locked memory, the number of the first parameter files is determined. Among them, the calculation execution order of the parameter files refers to the calculation sequence of the parameter files when the model performs forward propagation.
[0044] Based on the above embodiments of the present application, in the third embodiment of the present application, for the same or similar content as the above embodiments, reference can be made to the above introduction and will not be repeated hereinafter. On this basis, the step of reading the parameter file to be read from the expansion card to the buffer includes: determining the parameter file to be read in the expansion card based on the calculation execution order of the parameter file; if the parameter file to be read is not stored in the locked memory or the buffer, reading the parameter file to be read from the expansion card to the buffer.
[0045] Exemplarily, first, according to the calculation process of forward propagation, clarify the calculation execution order of each parameter file, and construct an ordered parameter file list. According to the current calculation time, determine the parameter file to be involved in the subsequent calculation from the above ordered parameter file list as the parameter file to be read. Check in turn whether the parameter file to be read is stored in the locked memory or the buffer. If the parameter file to be read is not stored in the locked memory or the buffer, read the parameter file to be read from the expansion card to the buffer, wherein the storage order of the parameter file read into the buffer can be consistent with the calculation execution order of the parameter file.
[0046] For example, as Figure 2 shown, if the target parameter file for forward propagation calculation at the current calculation time is parameter file 2, after determining that parameter file 2 is not stored in the locked memory or the buffer, first read parameter file 2 from the expansion card to the buffer, and then send parameter file 2 in the buffer to the video memory. At the same time, according to the calculation execution order of the parameter file, determine the parameter files to be read in the expansion card: parameter file 3 and parameter file 4. At this time, parameter file 3 has been stored in the buffer, so it does not need to be read into the buffer, while parameter file 4 is not stored in the buffer or the locked memory, so it is read from the expansion card to the buffer.
[0047] Optionally, the step of reading the parameter file to be read from the expansion card to the buffer includes: determining the number of parameter files to be read based on the remaining storage capacity of the buffer and the capacity of the parameter file to be read, and reading the parameter file to be read to the buffer based on the number of parameter files to be read and the calculation execution order of the parameter file.
[0048] Exemplarily, when reading the parameter file to be read from the expansion card into the buffer, the remaining storage capacity of the buffer is monitored in real time. The parameter files are arranged in the order of calculation execution, and the capacity of each parameter file is determined. Initialize a quantity variable to record the number of parameter files that can be read; initialize a capacity variable to record the remaining storage capacity of the buffer. According to the calculation execution order of the parameter files, starting from the first parameter file, each time a parameter file to be read is selected, the quantity variable is incremented by 1, and the capacity variable is subtracted by the capacity of the selected parameter file to be read. Repeat the above selection process until the capacity variable is less than the next parameter file to be selected. Then all the selected parameter files to be read are read from the expansion card into the buffer.
[0049] Based on the above embodiments of the present application, in the fourth embodiment of the present application, the same or similar content as the above embodiments can be referred to the above introduction and will not be repeated hereinafter. On this basis, after the forward propagation calculation at the current calculation moment is completed, it is determined whether the target parameter file is stored in the buffer. If the target parameter file is stored in the buffer, the target parameter file in the buffer is deleted.
[0050] Optionally, when performing the step of reading the parameter file from the expansion card into the buffer, if the remaining capacity of the buffer is not enough to store the parameter file, wait until the forward propagation calculation at the current calculation moment is completed, delete the target parameter file that has been executed from the buffer, and then read the parameter file from the expansion card into the buffer.
[0051] In this embodiment, by deleting the target parameter files that are no longer needed in the buffer, the memory space of the buffer can be released. This helps to avoid over-occupation of the buffer and ensure that there is enough space in the buffer to store the parameter files read from the expansion card subsequently.
[0052] Exemplarily, to help understand the implementation process of the calculation parameter acquisition method based on the DeepSpeed framework obtained by combining this embodiment with the above embodiments, please refer to Figure 4 , Figure 4A brief process schematic diagram of a method for obtaining calculation parameters based on the DeepSpeed framework is provided. Specifically: First, unload the first parameter file from the expansion card to the locked memory. The locked memory is pre-allocated and does not exchange data with other storage areas. At the current calculation moment, determine the target parameter file for forward propagation calculation, and check whether the target parameter file has been stored in the locked memory or the buffer. If the target parameter file is not in the locked memory or the buffer, read the target parameter file from the expansion card to the buffer; if the target parameter file is stored in the locked memory, read the target parameter file in the locked memory to the video memory; if the target parameter file is stored in the buffer, send the target parameter file in the buffer to the video memory. While reading the target parameter file in the buffer locked memory to the video memory, According to the calculation execution order of the parameter file, determine the parameter file to be read in the expansion card, and check whether the parameter file to be read has been stored in the locked memory or the buffer. If the parameter file to be read is not in the locked memory or the buffer, read the parameter file to be read from the expansion card to the buffer. Among them, after the forward propagation calculation at the current calculation moment is completed, determine whether the target parameter file is stored in the buffer. If the target parameter file is stored in the buffer, delete the target parameter file in the buffer.
[0053] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the method for obtaining calculation parameters based on the DeepSpeed framework of this application. Based on this technical concept, more forms of simple transformations are within the protection scope of this application.
[0054] This application provides a device for obtaining calculation parameters based on the DeepSpeed framework. The device for obtaining calculation parameters based on the DeepSpeed framework includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method for obtaining calculation parameters based on the DeepSpeed framework in the first embodiment above.
[0055] Next, refer to Figure 5 , which shows a schematic structural diagram of a device for obtaining calculation parameters based on the DeepSpeed framework suitable for implementing the embodiments of this application. The device for obtaining calculation parameters based on the DeepSpeed framework in the embodiments of this application may include, but is not limited to, mobile terminals such as laptop computers, tablet computers (PAD, Portable Application Description), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 5The illustrated computing parameter acquisition device based on the DeepSpeed framework is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.
[0056] As Figure 5 shown, the computing parameter acquisition device based on the DeepSpeed framework may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM, Read Only Memory) 1002 or the program read from the storage device 1003 into the random access memory (RAM, Random Access Memory) 1004. In the random access memory 1004, various programs and data required for the operation of the computing parameter acquisition device based on the DeepSpeed framework are also stored. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other through a bus 1005. The input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD, Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the computing parameter acquisition device based on the DeepSpeed framework to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows a computing parameter acquisition device based on the DeepSpeed framework with various systems, it should be understood that it is not required to implement or have all the shown systems. More or fewer systems may be alternatively implemented or had.
[0057] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program may be downloaded and installed from the network through the communication device, or installed from the storage device 1003, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above functions defined in the methods of the embodiments disclosed in the present application are executed.
[0058] The computing parameter acquisition device based on the DeepSpeed framework provided by this application adopts the computing parameter acquisition method based on the DeepSpeed framework in the above-mentioned embodiment, and can solve the technical problem of low computing efficiency of parameters during model training under the DeepSpeed framework. Compared with the prior art, the beneficial effects of the computing parameter acquisition device based on the DeepSpeed framework provided by this application are the same as those of the computing parameter acquisition method based on the DeepSpeed framework provided in the above-mentioned embodiment, and other technical features in the computing parameter acquisition device based on the DeepSpeed framework are the same as the features disclosed in the method of the previous embodiment, which will not be elaborated here.
[0059] It should be understood that each part disclosed in this application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.
[0060] As described above, this is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
[0061] This application provides a computer-readable storage medium with computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the computing parameter acquisition method based on the DeepSpeed framework in the above-mentioned embodiment.
[0062] The computer-readable storage medium provided by the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), or a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in conjunction with an instruction execution system, device, or component. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, radio frequency (RF), etc., or any suitable combination of the above.
[0063] The above computer-readable storage medium may be included in a computing parameter acquisition device based on the DeepSpeed framework; or it may exist separately without being assembled into a computing parameter acquisition device based on the DeepSpeed framework.
[0064] The above computer-readable storage medium carries one or more programs which, when executed by a computing parameter acquisition device based on the DeepSpeed framework, enable the computing parameter acquisition device based on the DeepSpeed framework to write computer program code for performing the operations of this application in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, or executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0065] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0066] The modules involved in the embodiments described in this application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation on the unit itself in some cases.
[0067] The readable storage medium provided by this application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned method for obtaining calculation parameters based on the DeepSpeed framework, and can solve the technical problem of low calculation efficiency of parameters during model training under the DeepSpeed framework. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by this application are the same as those of the method for obtaining calculation parameters based on the DeepSpeed framework provided in the above embodiments, and will not be elaborated herein.
[0068] The above are only some embodiments of this application, and thus do not limit the patent scope of this application. Any equivalent structural transformation made under the technical concept of this application by using the content of the specification and drawings of this application, or any direct / indirect application in other related technical fields, is included in the patent protection scope of this application.
Claims
1. A method for obtaining calculation parameters based on the DeepSpeed framework, characterized in that: At the nth calculation moment, the method comprises: Determine the target parameter file used for forward propagation calculation at the current calculation moment; Determining whether the target parameter file is stored in a locked memory or a cache area; If the target parameter file is stored in the buffer area, the target parameter file in the buffer area is sent to the video memory, and the parameter file is read from the expansion card to the buffer area.
2. The method for obtaining calculation parameters based on the DeepSpeed framework according to claim 1, characterized in that: Before the step of determining the target parameter file for forward propagation calculation at the current calculation moment, the method further includes: Determining the first parameter file to be unloaded to the locked memory and the quantity of the first parameter file based on the capacity of the locked memory and the capacity of all parameter files corresponding to the forward propagation calculation; Based on the number of the first parameter files, the first parameter files are unloaded from the expansion card to the locked memory.
3. The method for obtaining calculation parameters based on the DeepSpeed framework according to claim 1, characterized in that: After the step of determining whether the target parameter file is stored in the locked memory or the cache area, the method further includes: If the target parameter file is stored in the locked memory, the target parameter file in the locked memory is read into the video memory, and the parameter file is read from the expansion card into the buffer area.
4. The method for obtaining calculation parameters based on the DeepSpeed framework according to claim 1, characterized in that: After the step of determining whether the target parameter file is stored in the locked memory or the cache area, the method further includes: If the target parameter file is not stored in the locked memory or the cache area, after reading the target parameter file from the expansion card to the cache area, the steps of sending the target parameter file in the cache area to the video memory and reading the parameter file from the expansion card to the cache area are performed.
5. The method for obtaining calculation parameters based on the DeepSpeed framework according to any one of claims 1 to 4, characterized in that: The step of reading the parameter file from the expansion card to the buffer area comprises: Determining the parameter file to be read in the expansion card based on the calculation execution order of the parameter file; If the parameter file to be read is not stored in the locked memory or the cache area, the parameter file to be read is read from the expansion card to the cache area.
6. The method for obtaining calculation parameters based on the DeepSpeed framework according to claim 5, characterized in that: The step of reading the parameter file to be read from the expansion card to the buffer area comprises: Determining the number of parameter files to be read based on the remaining storage capacity of the cache area and the capacity of the parameter files to be read; Based on the number of the parameter files to be read and the calculation execution order of the parameter files, the parameter files to be read are read into the cache area.
7. The method for obtaining calculation parameters based on the DeepSpeed framework according to claim 1, characterized in that: The method further comprises: After the forward propagation calculation at the current calculation moment is completed, determining whether the target parameter file is stored in the buffer area; If the target parameter file is stored in the cache area, the target parameter file in the cache area is deleted.
8. The method for obtaining calculation parameters based on the DeepSpeed framework according to claim 2, characterized in that: The step of determining the first parameter file to be unloaded to the locked memory and the number of the first parameter files based on the capacity of the locked memory and the capacity of all parameter files corresponding to the forward propagation calculation comprises: Obtaining the historical calculation frequencies of all parameter files corresponding to the forward propagation calculation; Determining an uninstall weight value of the parameter file based on a historical calculation frequency and a capacity of the parameter file; The first parameter file and the number of the first parameter files are determined in all parameter files based on the uninstall weight value of the parameter file, the capacity of the locked memory and the capacity of the parameter file.
9. A computing parameter acquisition device based on the DeepSpeed framework, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the method for obtaining computing parameters based on the DeepSpeed framework as described in any one of claims 1 to 8.
10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the method for obtaining calculation parameters based on the DeepSpeed framework as described in any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
File reading method and device, electronic equipment and storage medium
CN110807010A
Distributed file caching method, system and equipment and storage medium
CN115269522A
Model parameter management method, host, equipment and storage medium
CN117669679A
Model optimization method and related device
CN118626415A
Training task check point file storage method and device, equipment and medium
CN118897824A
Cited By
Large language model efficient reasoning method and framework for resource-constrained end-side equipment
CN122021900A