Control method of training and pushing expansion card, training and pushing expansion card and storage medium

By using the control method of the training push expansion card on the home terminal, the video memory occupancy value is detected and the training push data is stored on the target expansion card, the problem of insufficient graphics card in the home terminal is solved and the low-cost local deployment of the large model is realized.

CN120144295APending Publication Date: 2025-06-13YEESTOR MICROELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510227584.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The graphics memory capacity of individual users’ home terminal graphics cards is insufficient and they cannot deploy large language models locally.

Method used

It provides a control method for the push-out expansion card. By detecting the push-out data output by the push-out unit, it determines the video memory occupancy value. If the threshold is exceeded, select the target push-out expansion card and store the push-out data to the target push-out expansion card through the direct connection bus.

Benefits of technology

It effectively reduces the video memory pressure of the training and pushing unit, realizes a low-cost local deployment model, and solves the problem of graphics card hardware limitations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144295A_ABST
    Figure CN120144295A_ABST
Patent Text Reader

Abstract

The invention discloses a training and pushing expansion card control method, a training and pushing expansion card and a storage medium, and relates to the technical field of electronic digital data processing, the training and pushing expansion card control method comprises the following steps: if it is detected that a training and pushing unit outputs training and pushing data, determining a video memory occupation value of the training and pushing data; if the video memory occupancy value is greater than a video memory threshold associated with the training and pushing unit, determining a target training and pushing expansion card; and storing the training and pushing data to the target training and pushing expansion card through a direct connection bus between the training and pushing unit and the target training and pushing expansion card. The technical problem that a large model cannot be trained and deployed locally due to the limitation of graphics card hardware is solved; a low-cost local deployment training large model is realized, and a common computer is changed into a machine supporting large model training and pushing by inserting a training and pushing expansion card.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of electronic digital data processing, and particularly to a control method for a training and inference expansion card, a training and inference expansion card, and a storage medium. Background Art

[0002] In the context of the increasing demand for current data privacy protection, local deployment of large language models has become an important option for users to avoid the risk of cloud data leakage. However, the video memory capacity of the graphics cards usually configured in personal user's home terminals is mostly 4GB - 16GB, which is far lower than the minimum video memory requirement for running large models. This serious imbalance between the video memory and the number of model parameters results in users being unable to locally train and deploy large models. Summary of the Invention

[0003] The main purpose of this application is to provide a control method for a training and inference expansion card, a training and inference expansion card, and a storage medium, aiming to solve the technical problem that large models cannot be locally trained and deployed due to the hardware limitations of graphics cards.

[0004] To achieve the above objective, this application provides a control method for a training and inference expansion card, and the control method for the training and inference expansion card includes:

[0005] If it is detected that the training and inference unit outputs training and inference data, determine the video memory occupancy value of the training and inference data;

[0006] If the video memory occupancy value is greater than the video memory threshold associated with the training and inference unit, determine the target training and inference expansion card;

[0007] Store the training and inference data to the target training and inference expansion card through the direct connection bus between the training and inference unit and the target training and inference expansion card.

[0008] In an embodiment, before the step of if it is detected that the training and inference unit outputs training and inference data, and determine the video memory occupancy value of the training and inference data, it includes:

[0009] If it is detected that at least one training and inference expansion card is connected, load the driver configuration of the training and inference expansion card;

[0010] Based on the driver configuration, the model of the training and inference unit, and the hardware resources of the terminal to which the training and inference expansion card is connected, configure a large model training and inference environment on the terminal.

[0011] In an embodiment, after the step of configure a large model training and inference environment on the terminal, it includes:

[0012] In response to a model training and inference instruction, determine the model identifier corresponding to the model training and inference instruction;

[0013] Determine the data scheduling policy corresponding to the model identifier according to the large model training and inference environment.

[0014] In one embodiment, after the step of determining a resource scheduling policy according to the model identifier and the configured large model training and inference environment, the following steps are included:

[0015] Obtain model parameters;

[0016] Based on the data scheduling policy, divide the model parameters into video memory data and expansion card data;

[0017] Store the expansion card data in the training and inference expansion card, and send the video memory data to the training and inference unit.

[0018] In one embodiment, before the step of dividing the model parameters into video memory data and expansion card data based on the data scheduling policy, the following steps are included:

[0019] In response to a video memory configuration instruction, determine the video memory quota of the training and inference expansion card corresponding to the video memory configuration instruction;

[0020] Update the data scheduling policy according to the video memory quota.

[0021] In one embodiment, the training and inference expansion card includes a processing core and storage particles, and a Linux system is burned in the processing core. The control method of the training and inference expansion card further includes:

[0022] In response to a main card configuration instruction, determine the first training and inference expansion card corresponding to the main card configuration instruction;

[0023] Run the Linux system of the first training and inference expansion card, and mask the Linux system of the second training and inference expansion card, where the first training and inference expansion card and the second training and inference expansion card form an expansion card array.

[0024] In one embodiment, the control method of the training and inference expansion card further includes:

[0025] In response to a pre-cache instruction, determine the target training and inference expansion card according to the pre-cache instruction;

[0026] Determine the pre-fetched data in the target training and inference expansion card;

[0027] Transmit the pre-fetched data to the training and inference unit through a direct connection bus.

[0028] In one embodiment, before the step of determining the video memory occupancy value of the training and inference data, the following steps are further included:

[0029] Determine whether the training and inference data is used for the next frame calculation;

[0030] If it is not used for the next frame calculation, execute the step of determining the video memory occupancy value of the training and inference data.

[0031] In addition, to achieve the above object, the present application further provides a training and inference expansion card, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the control method of the training and inference expansion card as described above.

[0032] In addition, to achieve the above object, the present application further provides a storage medium, which is a computer-readable storage medium, and a program for implementing the control method of the training and inference expansion card is stored on the computer-readable storage medium, and the program for implementing the control method of the training and inference expansion card is executed by a processor to implement the steps of the control method of the training and inference expansion card as described above.

[0033] The present application provides a control method for a training and inference expansion card. First, the present application determines the video memory occupancy value of the training and inference data by detecting that the training and inference unit outputs the training and inference data; if the video memory occupancy value is greater than the video memory threshold associated with the training and inference unit, determines the target training and inference expansion card; stores the training and inference data to the target training and inference expansion card through a direct connection bus between the training and inference unit and the target training and inference expansion card. The technical problem that the hardware limitation of the graphics card causes the inability to locally train and deploy large models is solved; the low-cost local deployment of large models is realized. Description of the Drawings

[0034] The drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0035] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0036] Figure 1 It is a schematic flowchart provided for Embodiment 1 of the control method of the training and inference expansion card of the present application;

[0037] Figure 2 It is a schematic flowchart provided for Embodiment 2 of the control method of the training and inference expansion card of the present application;

[0038] Figure 3 It is a schematic flowchart provided for Embodiment 3 of the control method of the training and inference expansion card of the present application;

[0039] Figure 4 It is a schematic hardware structure diagram of the training and inference expansion card of the present application;

[0040] Figure 5 This is another schematic diagram of the hardware structure of the training and inference expansion card of this application.

[0041] The realization of the purpose, functional characteristics and advantages of this application will be further described with reference to the accompanying drawings in combination with embodiments. Detailed implementation manners

[0042] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of this application and are not used to limit this application.

[0043] In order to better understand the technical solutions of this application, the following will be described in detail in combination with the drawings of the specification and specific implementation manners.

[0044] Currently, in the context of the increasing demand for current data privacy protection, local deployment of large language models has become an important choice for users to avoid the risk of cloud data leakage. However, the video memory capacity of the graphics cards usually configured in personal user's home terminals is mostly 4GB - 16GB, which is much lower than the minimum video memory requirement for running large models. This serious imbalance between the video memory and the number of model parameters results in users being unable to locally train and deploy large models.

[0045] The main solution of this application is: if it is detected that the training and inference unit outputs training and inference data, determine the video memory occupancy value of the training and inference data; if the video memory occupancy value is greater than the video memory threshold associated with the training and inference unit, determine the target training and inference expansion card; store the training and inference data to the target training and inference expansion card through the direct connection bus between the training and inference unit and the target training and inference expansion card.

[0046] This application realizes the low-cost local deployment of large models by connecting the training and inference expansion card to the terminal, then interacting with the training and inference unit of the terminal through a direct connection bus, and combining with the dynamic scheduling of training and inference data.

[0047] It should be noted that the execution subject of this embodiment can be the training and inference expansion card, or a computing service device with data processing, network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc. This embodiment does not make specific limitations on this.

[0048] Based on this, Embodiment 1 of this application proposes a control method for a training and inference expansion card. Please refer to Figure 1 , the control method of the training and inference expansion card includes steps S10 to S30:

[0049] Step S10, if it is detected that the training and inference unit outputs training and inference data, determine the video memory occupancy value of the training and inference data.

[0050] It should be noted that the training and inference expansion card in this application includes a processing core and memory particles. The linux system is burned into the processing core, which can directly execute the control method of the training and inference expansion card in this application to complete the interaction with the training and inference unit. It can also be executed by the processor of the terminal connected to the training and inference expansion card to complete the interaction between the training and inference expansion card and the training and inference unit. It can also be that the processor of the terminal is the main system and the linux system of the training and inference expansion card is the secondary system to complete the interaction. Specifically, when the two systems coexist, the main system is used to access data from the training and inference expansion card, and the secondary system is used to schedule data between the training and inference expansion card arrays.

[0051] In this embodiment, the training and inference unit: the device that undertakes the main computing tasks during the deep learning training and inference processes, generally a high-performance GPU (graphics processing unit), such as NVIDIA's A100, V100 and other graphics cards. They have powerful parallel computing capabilities and can quickly process large-scale matrix operations, and are the core hardware for deep learning model training and inference. Training and inference data: the data generated during the model training and inference phases, including the input training samples, intermediate activation values of the model, gradient data, as well as model parameters, optimizer states, checkpoints, etc. For example, in the training of an image recognition model, the training and inference data can be image data, the feature maps output by the convolutional layer, and the gradients obtained by backpropagation. Video memory occupancy value: the size of the storage space occupied by the training and inference data in the video memory of the training and inference unit, measured in bytes. Video memory is the high-speed memory used by the GPU to temporarily store data, and the video memory occupancy value reflects the consumption of video memory resources by the training and inference data.

[0052] As an alternative implementation, a data monitoring mechanism is set at the output port of the training and inference unit, and the data detection function can be implemented through a hardware circuit or a software driver program. For example, a flag bit is set in the output register of the GPU. When new training and inference data is generated, the flag bit is set, thereby triggering the data detection program. When the training and inference data output is detected, the data type of the training and inference data needs to be identified first. Common data types include integers, floating-point numbers, boolean values, etc. Different data types occupy different storage spaces in the video memory. For example, a 32-bit floating-point number usually occupies 4 bytes of memory space. For multi-dimensional data (such as tensors), the sizes of its various dimensions need to be calculated. By multiplying the number of bytes occupied by the data type by the total number of elements of the data (i.e., the product of the sizes of each dimension), the video memory occupancy value of the training and inference data can be obtained. For example, a 32-bit floating-point number tensor with a shape of (3, 4, 5) has a total number of elements of 3×4×5 = 60, and each element occupies 4 bytes, so the video memory occupancy value of this tensor is 60×4 = 240 bytes.

[0053] Step S20: If the video memory occupancy value is greater than the video memory threshold associated with the training and inference unit, determine the target training and inference expansion card.

[0054] In this embodiment, the video memory threshold is an upper limit value set for the video memory occupancy of the training and inference data. When the video memory occupancy value of the training and inference data exceeds this threshold, it means that the video memory usage of the training and inference unit is close to or reaches saturation. At this time, some training and inference data need to be transferred to other storage devices to avoid training or inference interruption caused by video memory overflow.

[0055] The target training and inference expansion card is a specific expansion card selected from multiple training and inference expansion cards for storing training and inference data. The training and inference expansion card is connected to the training and inference unit through a high-speed interface (such as PCIe), has a large storage capacity, and can relieve the video memory pressure of the training and inference unit.

[0056] As an alternative implementation, compare the calculated video memory occupancy value of the training and inference data with the video memory threshold associated with the training and inference unit. The video memory threshold can be set according to the video memory capacity of the training and inference unit and the actual application requirements. For example, for a GPU with a video memory capacity of 8GB, the video memory threshold can be set to 6GB. When the video memory occupancy value is greater than the video memory threshold, collect the relevant information of all training and inference expansion cards, including the remaining storage capacity, read and write speed, health status, etc. of the expansion cards. These information can be obtained by querying the management registers of the expansion cards or calling the driver program interfaces of the expansion cards. According to the collected expansion card information, adopt a certain selection strategy to determine the target training and inference expansion card. Common selection strategies include selecting the expansion card with the largest remaining storage capacity, selecting the expansion card with the fastest read and write speed, etc. For example, select the expansion card with the largest remaining storage capacity as the target expansion card to ensure that there is enough space to store the training and inference data.

[0057] As another alternative implementation, obtain the usage conditions of each training and inference expansion card, obtain the preset data scheduling strategy, and determine the target training and inference expansion card according to the data scheduling strategy and the obtained usage conditions of each training and inference expansion card.

[0058] Specifically, the data scheduling strategy is a pre-set rule and method for reasonably allocating and scheduling training and inference data among multiple training and inference expansion cards. Its purpose is to make full use of the performance and resources of each training and inference expansion card, improve the efficiency and performance of the entire training and inference system, and avoid the situation where some expansion cards are overused while some expansion cards are idle. The data scheduling strategy can comprehensively consider multiple factors of the training and inference expansion cards, such as remaining storage capacity, read and write speed, computing power, load conditions, etc., to achieve optimal data allocation and processing.

[0059] The usage of the training and inference expansion card includes multiple key metrics, and the methods to obtain these metrics are as follows: Remaining storage capacity: By communicating with the training and inference expansion card, query the status information of its storage device. For example, for a flash-based training and inference expansion card, a specific storage management protocol or driver interface can be used to obtain the remaining available storage space. In the Linux system, relevant system commands can be called or programs can be written to read the storage information of the expansion card. Read and write speed: By writing and reading a certain size of data blocks to the training and inference expansion card, record the time required for the operation, and then calculate the read and write speed. For example, using a standard file read and write test tool, create a test file of a specific size on the expansion card, and record the time for writing and reading the file respectively, so as to obtain the numerical value of the read and write speed. Computational power: Evaluate the computational power of the training and inference expansion card by running some benchmark test programs. For example, using the benchmark test tools provided by deep learning frameworks, test the expansion card for common deep learning computational tasks such as matrix operations and convolution operations, and evaluate its computational power according to the test results. Load condition: Real-time monitor the number of tasks being processed by the training and inference expansion card, the complexity of the tasks, and the processing progress and other information to determine its load condition. These information can be obtained through the monitoring interface of the expansion card or the system log.

[0060] The data scheduling strategy is usually set during the system initialization phase and stored in the system configuration file or database. Determine the storage location of the data scheduling strategy, which may be a local configuration file (such as JSON, XML format) or a remote database. According to the different storage locations, adopt corresponding methods to read the data scheduling strategy. If it is a local configuration file, file reading functions can be used to read the file content; if it is a remote database, database connection tools and query statements can be used to obtain the strategy information.

[0061] Optionally, preferentially select the training and inference expansion card with the largest remaining storage capacity as the target expansion card to ensure sufficient space for storing training and inference data. Sort the remaining storage capacities of the obtained training and inference expansion cards. Select the training and inference expansion card with the largest remaining storage capacity as the target expansion card.

[0062] Optionally, preferentially select the training and inference expansion card with the fastest read and write speed as the target expansion card to improve the efficiency of data transmission and processing. Sort the read and write speeds of the obtained training and inference expansion cards. Select the training and inference expansion card with the fastest read and write speed as the target expansion card.

[0063] Optionally, considering multiple factors such as the remaining storage capacity, read / write speed, computing power, and load condition of the training and inference expansion card, calculate a comprehensive score for each expansion card, and select the expansion card with the highest score as the target expansion card. Assign a weight to each evaluation factor. For example, the weight of the remaining storage capacity is 0.3, the weight of the read / write speed is 0.3, the weight of the computing power is 0.2, and the weight of the load condition is 0.2. Standardize each evaluation factor of each training and inference expansion card so that its value range is between 0 and 1. Calculate the comprehensive score of each training and inference expansion card according to the weight and the standardized evaluation factor value. The calculation formula is: Comprehensive score = remaining storage capacity score × 0.3 + read / write speed score × 0.3 + computing power score × 0.2 + load condition score × 0.2. Select the training and inference expansion card with the highest comprehensive score as the target expansion card.

[0064] Step S30, store the training and inference data to the target training and inference expansion card through the direct connection bus between the training and inference unit and the target training and inference expansion card.

[0065] In this embodiment, the direct connection bus: a high-speed data transmission channel directly connected between the training and inference unit and the training and inference expansion card, such as a PCIe bus. The direct connection bus provides low-latency and high-bandwidth data transmission capabilities, and can quickly transmit the training and inference data from the training and inference unit to the target training and inference expansion card. The training and inference expansion card is connected to the PCIe bus through PCIe P2P so that the training and inference expansion card can directly interact with the training and inference unit for data.

[0066] Furthermore, the training and inference expansion cards interact through CXL (Compute Express Link), and then perform storage pooling, that is, each training and inference expansion card selects a part of its capacity according to its own capacity to form a common storage unit. The common storage unit interacts with the training and inference unit through the PCIe bus for data.

[0067] As an alternative implementation, the training and inference expansion card includes a buffer area and a storage area, and the read / write speed of the buffer area is greater than that of the storage area. Each training and inference expansion card performs storage pooling by interacting the buffer areas through CXL, so that when writing training and inference data, if the buffer area capacity of the target training and inference expansion card is insufficient, the training and inference data is transmitted to the common storage unit in batches through CXL. Then, the background transfers the training and inference data from the buffer area to the storage area to improve the writing speed of the training and inference data.

[0068] Furthermore, if the training and inference unit supports CXL, each training and inference expansion card is directly connected to the training and inference unit through CXL, or each training and inference expansion card performs storage pooling through CXL, and then the common storage unit is directly connected to the training and inference unit through CXL.

[0069] As an alternative implementation, before data transmission, some preparatory work needs to be done, such as configuring the transmission parameters of the direct connection bus (such as transmission rate, data block size, etc.), initializing the data transmission buffer, etc. These configuration tasks can be accomplished by setting the registers of the bus controller. Start the data transmission process, and transfer the training and inference data from the video memory of the training and inference unit to the storage medium of the target training and inference expansion card through the direct connection bus. During the transmission process, DMA (Direct Memory Access) technology can be adopted to enable the data to be directly transmitted between the training and inference unit and the target training and inference expansion card without CPU intervention, thereby improving the data transmission efficiency. After the data transmission is completed, the target training and inference expansion card sends a transmission confirmation message to the training and inference unit, indicating that the data has been successfully stored. After receiving the confirmation message, the training and inference unit releases the relevant video memory resources.

[0070] Exemplarily, use the TensorFlow framework to train a language model based on the Transformer architecture. The training and inference unit is an NVIDIA RTX 4090 GPU with a video memory capacity of 24GB, and the video memory threshold is set to 100MB. There are 4 training and inference expansion cards connected in the system, namely expansion card X, expansion card Y, expansion card Z, and expansion card W, and their remaining storage capacities are 129GB, 73GB, 529GB, and 326GB respectively. During the training process, the training and inference unit (RTX 4090 GPU) performs forward propagation and backward propagation calculations, generating intermediate activation values and gradient data as training and inference data. When the gradient data output is detected, identify its data type as 16-bit floating-point number and calculate its dimension size. Assuming the shape of the gradient tensor is (32, 64, 128, 256), then the total number of its elements is 32×64×128×256 = 67108864, and each element occupies 2 bytes, so the video memory occupancy value of this gradient data is 67108864×2 = 134217728 bytes (about 128MB). The video memory usage will exceed the video memory threshold of 100MB. Therefore, it is necessary to store the gradient data in the training and inference expansion card. Collect the remaining storage capacity information of the 4 training and inference expansion cards and find that the remaining storage capacity of expansion card Z is the largest, which is 529GB. So, select expansion card Z as the target training and inference expansion card. Configure the transmission parameters of the PCIe bus, set the transmission rate to the highest, and initialize the data transmission buffer. Adopt DMA technology to quickly transfer the gradient data from the video memory of the training and inference unit to the storage medium of expansion card Z through the PCIe bus.

[0071] As another alternative implementation, if the video memory occupancy value is greater than the video memory threshold associated with the training and inference unit, determine the target training and inference expansion card, where the target training and inference expansion card is all the training and inference expansion cards, determine the number of training and inference expansion cards, divide the training and inference data into the corresponding number of parts according to the number of training and inference expansion cards, and store each part of the training and inference data in the corresponding training and inference expansion card.

[0072] Exemplarily, 6 training and inference expansion cards are connected in the system, namely expansion card 1, expansion card 2, expansion card 3, expansion card 4, expansion card 5 and expansion card 6, and their remaining storage capacities are 178GB, 73GB, 529GB, 326GB, 138GB and 69GB respectively. During the training process, the training and inference unit performs forward propagation and backward propagation calculations, generating intermediate activation values and gradient data as training and inference data. The video memory occupancy value of the training and inference data is 4.8GB. At this time, six training and inference expansion cards are determined, and then the training and inference data is divided into six parts. The division method can be equal division or proportional distribution according to the remaining storage capacity of each training and inference expansion card; after determining the six parts of training and inference sub-data, each part of the training and inference sub-data is stored in the corresponding training and inference expansion card.

[0073] At this time, by forming a raid0 with each training and inference expansion card and then dividing the training and inference data into n parts, where n is the number of training and inference expansion cards, the technical effect of increasing the transmission speed of the training and inference data to n times the original is achieved.

[0074] Through this embodiment, according to the video memory occupancy value and video memory threshold of the training and inference data, a suitable target training and inference expansion card is selected, and the training and inference data is stored in the expansion card through a direct connection bus, thereby effectively reducing the video memory pressure of the training and inference unit and ensuring the smooth progress of model training.

[0075] Based on the above embodiment, Embodiment 2 of the present application proposes a control method for a training and inference expansion card. Refer to Figure 2 , before step S10, it includes:

[0076] Step A10, if at least one training and inference expansion card is detected to be connected, load the driver configuration of the training and inference expansion card.

[0077] In this embodiment, the driver configuration: is a set of software programs and setting information for enabling the terminal operating system to recognize, control and communicate with the training and inference expansion card. The driver configuration contains information such as the hardware parameters, functional characteristics of the expansion card, and interfaces for interacting with the operating system, ensuring that the expansion card can work properly.

[0078] The hardware management module of the terminal continuously monitors the status of specific interfaces (such as PCIe slots). When a training and inference expansion card is inserted into the interface, the hardware management module will detect changes in interface levels, signals, etc., thereby determining that a training and inference expansion card has been connected. For example, the PCIe interface will detect the electrical connection of the new device, triggering the device detection mechanism of the system. The terminal operating system searches for a matching driver configuration file in the local driver library or the remote driver server in the networked state based on the detected hardware identification information of the training and inference expansion card (such as vendor ID, device ID). The driver library is usually stored in a specific directory of the operating system and contains driver programs for various hardware devices. Once a suitable driver configuration file is found, the operating system loads it into memory and performs initialization operations. The initialization process includes allocating system resources, setting hardware registers, establishing a communication channel with the expansion card, etc., enabling the training and inference expansion card to perform normal data interaction with the terminal system.

[0079] Step A20, configure the large model training and inference environment on the terminal based on the driver configuration, the model of the training and inference unit, and the hardware resources of the training and inference expansion card accessing the terminal.

[0080] In this embodiment, different models of training and inference units have differences in computing performance, video memory capacity, power consumption, etc., and these factors will affect the efficiency and effect of large model training and inference. Hardware resources: various hardware components owned by the terminal and their performance indicators, including the number of cores and main frequency of the CPU (Central Processing Unit), memory capacity, read and write speeds of storage devices, etc. The status of the hardware resources will limit the scale and speed of large model training and inference, and it needs to be fully considered when configuring the training and inference environment. Large model training and inference environment: a comprehensive environment including software and hardware settings, used to support the training and inference of large-scale deep learning models. It includes the installation and configuration of the operating system, deep learning frameworks (such as TensorFlow, PyTorch), computing libraries such as CUDA, as well as the allocation and scheduling of hardware resources.

[0081] Integrate the loaded training and inference expansion card driver configuration with the terminal system to ensure that the operating system can correctly recognize and call the functions of the expansion card. For example, in the Linux system, by modifying the kernel module and device driver parameters, the system can manage the expansion card as an available hardware device. Install a compatible deep learning framework and computing library according to the model of the training and inference unit. Different models of GPUs may require different versions of CUDA drivers and deep learning frameworks to achieve the best performance. For example, newer NVIDIA GPUs may require higher versions of CUDA to support their new computing features. At the same time, configure the framework so that it can utilize the computing resources of the training and inference unit and the training and inference expansion card. Conduct a comprehensive evaluation of the hardware resources accessed by the training and inference expansion card on the terminal, including CPU performance, memory capacity, storage speed, etc. According to the evaluation results, reasonably allocate hardware resources for large model training and inference. For example, if the memory capacity is limited, memory optimization strategies such as batch data loading and model parameter compression can be adopted. Set various environmental parameters required for large model training and inference on the terminal, such as learning rate, batch size, number of training epochs, etc. These parameters will affect the training effect and convergence speed of the model and need to be adjusted according to the specific model and task. At the same time, configure relevant settings such as data storage paths and logging to facilitate monitoring and management of the training and inference process.

[0082] Exemplarily, training and inference of a large language model with 500B parameters are performed on a personal PC. The PC was originally equipped with an NVIDIA RTX 3080 GPU as the training and inference unit. The user newly purchased two training and inference expansion cards and connected them to the PC through the PCIe interface. The user inserted the two training and inference expansion cards into the PCIe slots of the PC. The hardware management module of the PC detected the change in the interface status and determined that there were training and inference expansion cards connected. According to the hardware identification information of the expansion cards, the operating system found the matching driver configuration file in the local driver library and loaded it into the memory for initialization. After the initialization was completed, the system successfully recognized and was able to communicate with the training and inference expansion cards. The operating system integrated the driver configuration of the training and inference expansion cards with the system, and the newly connected training and inference expansion card devices could be seen in the system device manager. Since the training and inference unit is an NVIDIA RTX 3080 GPU, the user installed the compatible CUDA 11.3 driver and PyTorch 1.10 deep learning framework. Configure PyTorch to be able to utilize the computing resources of both the RTX 3080 and the training and inference expansion cards simultaneously. Evaluate the hardware resources of the PC. The CPU is an Intel Core i9-12900K with 16 cores and 24 threads, the memory capacity is 64GB, and the storage is a 1TB NVMe SSD. According to these resource conditions, a strategy of loading data in batches is determined to avoid memory overflow. Set the environment parameters for training and inference of the large language model in PyTorch. The learning rate is set to 0.0001, the batch size is set to 16, and the number of training epochs is set to 10. At the same time, configure the data storage path to a specific directory on the SSD for fast data reading and writing, and set up a log file to monitor metrics such as loss value and accuracy during the training and inference process.

[0083] Through the above steps, the large model training and inference environment was successfully configured on the terminal.

[0084] Optionally, after step A20, it includes:

[0085] Step A30, in response to the model training and inference instruction, determine the model identifier corresponding to the model training and inference instruction.

[0086] In this embodiment, the model training and inference instruction: a command issued by the user or the system to trigger the large model to perform training or inference operations. It contains the basic information for executing the training and inference task, such as the model to be used, the specific requirements for training and inference, etc. The model identifier: a unique identifier assigned to each large model, used to distinguish different large models. Through the model identifier, the corresponding large model can be accurately located and called for training and inference operations.

[0087] As an alternative implementation, the terminal system monitors the input of the model training and inference instructions in real time. When an instruction is received, it is parsed to extract the key information. The model training and inference instructions can be input in various ways, such as command-line input, graphical interface operations, etc. The parsing process converts the instruction string into a data structure that the system can understand. In the parsed instruction data structure, a field related to the model identifier is searched for. This field is usually predefined to clearly specify the large model to be used. For example, the instruction may contain information such as "model_id:123", where "123" is the corresponding model identifier.

[0088] Step A40: Determine the data scheduling strategy corresponding to the model identifier according to the large model training and inference environment.

[0089] In this embodiment, the data scheduling strategy is a set of rules and methods predefined for a specific large model to reasonably allocate and schedule training and inference data. It takes into account various factors in the training and inference environment, such as hardware performance, data characteristics, etc., to optimize the efficiency and performance of the training and inference process.

[0090] As an alternative implementation, comprehensively review the previously configured large model training and inference environment and collect the key information therein. This information includes the model and performance parameters of the training and inference units, the number of training and inference expansion cards, storage capacity and read / write speed, the type and version of the operating system, the configuration of the deep learning framework, etc. Establish a mapping relationship table between the model identifier and the data scheduling strategy. This table can be stored in the system configuration file or database. According to the determined model identifier, search for the corresponding data scheduling strategy in the mapping relationship table. For example, if the model identifier is "123", the mapping relationship table may record that this model corresponds to the "Comprehensive Performance Evaluation Strategy". For example, in the comprehensive performance evaluation strategy, different training and inference environments may need to adjust the weights of various evaluation factors such as remaining storage capacity, read / write speed, etc. to ensure that the strategy can better adapt to the actual situation.

[0091] Through the above steps, the system determines the model identifier according to the model training and inference instructions, and determines a suitable data scheduling strategy for this model according to the large model training and inference environment to improve the efficiency of local deployment and training of the large model.

[0092] Optionally, after step A40, it includes:

[0093] Step A50: Obtain model parameters.

[0094] In this embodiment, model parameters: a set of numerical values learned by the large model during training, and these values determine the behavior and performance of the model. For example, in a neural network model, model parameters include connection weights between neurons, bias terms, etc. They are the core components of the model and are used to calculate and predict input data.

[0095] As an alternative implementation, the model parameters may be stored in a specific folder on the local disk, or may be stored in a remote server or cloud storage. The storage location can be specified through a configuration file, environment variables, or hard-coded means. According to the determined storage location, the corresponding reading method is used to obtain the model parameters. If it is local storage, a file reading function can be used; if it is remote storage, network requests and corresponding protocols (such as HTTP, FTP) are required to download the parameters. For example, in Python, the torch.load() function is used to read the parameters of a PyTorch model.

[0096] Step A60, split the model parameters into video memory data and expansion card data based on the data scheduling strategy.

[0097] In this embodiment, video memory data: the data part split from the model parameters that is suitable for storage in the video memory of the training and inference unit. The video memory has the characteristics of high-speed reading and writing. Placing frequently used data in the video memory can improve the speed of model training and inference. Expansion card data: the data part split from the model parameters that needs to be stored in the training and inference expansion card. When the model parameters are too large to be fully accommodated in the video memory of the training and inference unit, part of the data is stored in the training and inference expansion card to relieve the video memory pressure.

[0098] As an alternative implementation, according to the previously determined data scheduling strategy, such as the remaining storage capacity priority strategy, the reading and writing speed priority strategy, or the comprehensive performance evaluation strategy, etc., the model parameters are evaluated. For example, under the remaining storage capacity priority strategy, the remaining space of the video memory of the training and inference unit and the remaining storage capacity of the training and inference expansion card will be given priority consideration. According to the evaluation results, the model parameters are split into video memory data and expansion card data. The split can be made according to factors such as the usage frequency and importance of the parameters. For example, for parameters that are frequently accessed, they are used as video memory data; for parameters that are not frequently used, they are used as expansion card data. It can also be split according to a certain ratio, such as taking the first 10% of the parameters as video memory data and the latter 90% of the parameters as expansion card data.

[0099] Step A70, store the expansion card data in the training and inference expansion card, and send the video memory data to the training and inference unit.

[0100] In this embodiment, ensure that the training and inference expansion card is correctly connected to and initialized by the terminal system, and establish a communication channel through the corresponding interface protocol (such as the PCIe protocol). Use the P2PDMA (Direct Memory Access) technology or other efficient data transfer methods to transfer the expansion card data from the system memory to the storage medium of the training and inference expansion card. During the transfer process, it is necessary to monitor the transfer status to ensure that the data is stored accurately without errors.

[0101] Further, before sending the video memory data, check whether there is enough space in the video memory of the training and inference unit to store this data. If the video memory is insufficient, first release some unnecessary data or adjust the data segmentation strategy.

[0102] Optionally, the video memory configuration in the data scheduling strategy can be modified. Before or after determining the data scheduling strategy corresponding to the model identifier according to the large model training and inference environment, the user can operate the video memory quota of each training and inference expansion card in the BIOS or driver. In response to the video memory configuration instruction, determine the video memory quota corresponding to the video memory configuration instruction; update the data scheduling strategy according to the video memory quota.

[0103] In this embodiment, the video memory configuration instruction: an instruction issued by the user or the system for configuring the use of the video memory of the training and inference expansion card. This instruction usually contains the setting requirements for the video memory quota and can be input through command lines, graphical interfaces, or system configuration files, etc. The video memory quota: the upper limit value of the available video memory allocated to the training and inference expansion card. It stipulates the maximum video memory space that the training and inference expansion card can use during the training and inference process to reasonably manage and allocate the system's video memory resources.

[0104] As an optional implementation method, the system listens for the input of the video memory configuration instruction in real time. When the instruction is received, it is parsed to extract the information related to the training and inference expansion card and the video memory quota. For example, the instruction may be "set_memory_quota--card_id2--quota195GB", where "--card_id2" indicates that the number of the training and inference expansion card to be configured is 2, and "--quota195GB" indicates that the allocated video memory quota is 195GB. According to the identified training and inference expansion card identifier (such as number, model, etc.) obtained by parsing, identify the corresponding training and inference expansion card in the system. The identification can be completed by querying the system's device list or using a device management tool. Extract the clear video memory quota value from the instruction and use it as the video memory quota of the training and inference expansion card. For example, determine 195GB in the above instruction as the video memory quota of the training and inference expansion card numbered 2.

[0105] In this embodiment, by detecting the access of the training and inference expansion card, the driver configuration is automatically loaded for the terminal when the training and inference expansion card is accessed. Then, based on the hardware resources of the terminal and the model type of the training and inference unit, the large model training and inference environment is configured. Furthermore, when the user deploys the large model, the system automatically scans the hardware resources and formulates an optimal resource matching plan for the specific hardware configuration.

[0106] Based on any of the above embodiments, Embodiment 3 of the present application proposes a control method for a training and inference expansion card. Referring to Figure 3 , it further includes:

[0107] Step B10, in response to the main card configuration instruction, determine the first training and inference expansion card corresponding to the main card configuration instruction.

[0108] In this embodiment, the main card configuration instruction: an instruction issued by the user or the system to specify a certain expansion card in the training and inference expansion card array as the main card. This instruction can be input through command lines, graphical user interfaces, or system configuration files, etc. The purpose is to clarify which expansion card in the expansion card array will undertake the main control and coordination tasks. The first training and inference expansion card: the training and inference expansion card designated as the main card according to the main card configuration instruction in the training and inference expansion card array. The main card will play a core role in the operation of the entire expansion card array, responsible for communicating with other expansion cards (the second training and inference expansion cards), coordinating data transmission, and task allocation, etc. The training and inference expansion card includes a processing core and storage particles, and the linux system is burned in the processing core.

[0109] As an optional implementation, the system listens to the input of the main card configuration instruction in real time. When the instruction is received, it is parsed to extract the identification information related to the training and inference expansion card to be designated as the main card. For example, the instruction may be "set_master_card--card_id1", where "--card_id1" indicates that the training and inference expansion card numbered 1 is to be designated as the main card. According to the parsed training and inference expansion card identification, the corresponding training and inference expansion card is searched in the device management list of the system. The system will check whether the expansion card exists in the expansion card array and whether its status is normal. The expansion card can be accurately identified by querying information such as the hardware ID and serial number of the expansion card. The identified training and inference expansion card that meets the instruction requirements is determined as the first training and inference expansion card (main card).

[0110] Step B20, run the linux system of the first training and inference expansion card, and mask the linux system of the second training and inference expansion card, where the first training and inference expansion card and the second training and inference expansion card form an expansion card array.

[0111] In this embodiment, the Linux system: an operating system running in the processing core of the training and inference expansion card, which provides support for the hardware resource management, task scheduling, and data processing of the expansion card. The Linux systems on different training and inference expansion cards can run independently, but in the expansion card array, they need to be coordinated according to the configuration of the main card. The second training and inference expansion card: in the expansion card array of training and inference expansion cards, other training and inference expansion cards except the first training and inference expansion card (main card). These expansion cards will work under the coordination of the main card, and the running state of their Linux systems will be adjusted according to the main card configuration. The expansion card array: a collection composed of multiple training and inference expansion cards connected and organized in a certain way, aiming to improve the overall performance and storage capacity of the training and inference system through collaborative work.

[0112] The system sends a startup signal to the first training and inference expansion card to trigger the startup of the Linux system in its processing core. During the startup process, the Linux system will perform a series of initialization operations, including loading driver programs, allocating system resources, starting system services, etc. To avoid conflicts and interference between multiple Linux systems, the system will take measures to shield the Linux system of the second training and inference expansion card. The specific shielding method can be through hardware switches, software instructions, or modifying system configurations, etc. For example, by sending a specific control signal to the second training and inference expansion card to make its Linux system enter the sleep state or prohibit it from starting some key services. After completing the startup of the Linux system of the first training and inference expansion card and the shielding operation of the Linux system of the second training and inference expansion card, the system will continuously monitor the status of each expansion card.

[0113] In this embodiment, by configuring and running the Linux system of the first training and inference expansion card, the system latency is further reduced, so that the processor of the terminal does not need to participate in training and inference. Even if the storage medium of the terminal is removed or damaged, training and inference can still continue to be executed.

[0114] Based on the above embodiments, Embodiment 4 of this application proposes a control method for a training and inference expansion card, which further includes:

[0115] Step C10, in response to a pre-cache instruction, determine the target training and inference expansion card according to the pre-cache instruction.

[0116] In this embodiment, the pre-cache instruction: an instruction issued by the system or the user to trigger the pre-cache operation. During the training and inference process of the large model, in order to improve the efficiency of data processing, some data is pre-cached from the training and inference expansion card to the training and inference unit in advance, and the pre-cache instruction is used to start this operation.

[0117] As an alternative implementation, the system monitors the input of the pre-cache instruction in real time. When the instruction is received, it is parsed to extract the key information related to the target training and inference expansion card, such as the identifier (number, model, etc.) of the training and inference expansion card. The instruction may be input in the form of a command line, such as "precache--card_id2", indicating that the training and inference expansion card numbered 2 is to be selected as the target training and inference expansion card. Based on the identifier of the training and inference expansion card obtained by parsing, the corresponding training and inference expansion card is searched for in the device list of the system.

[0118] Step C20: Determine the pre-fetched data in the target training and inference expansion card.

[0119] In this embodiment, the pre-fetched data is the data selected from the target training and inference expansion card and prepared to be transferred to the training and inference unit in advance. These data are usually those to be used during the training and inference process, such as model parameters, training samples, etc. Caching them in the training and inference unit in advance can reduce the waiting time for data reading and improve the training and inference efficiency.

[0120] As an alternative implementation, according to the data access rules specified in the pre-cache instruction, determine which data to pre-fetch from the target training and inference expansion card. These rules can be formulated based on factors such as the usage frequency of the data, access order, data dependency, etc. For example, according to the usage frequency of the data, pre-fetch the frequently used data first. Locate the data that conforms to the pre-fetch rule in the storage granules of the target training and inference expansion card. The data can be quickly located through methods such as indexing and tagging. At the same time, filter the located data to remove the unnecessary parts to ensure that only the data related to the current training and inference task is pre-fetched.

[0121] Step C30: Transmit the pre-fetched data to the training and inference unit through a direct connection bus.

[0122] In this embodiment, through the interface of the direct connection bus, start the transmission process of the pre-fetched data. Adopt the DMA (Direct Memory Access) technology to directly transfer the data between the target training and inference expansion card and the training and inference unit, reducing the intervention of the CPU and improving the transmission speed.

[0123] Exemplarily, in a deep learning-based image recognition model training and inference system, when the system receives a pre-cache instruction, it needs to prefetch some model parameters from the training and inference expansion card to the training and inference unit to accelerate the subsequent inference speed. There are 3 training and inference expansion cards in the system, numbered 1, 2, and 3 respectively. The system receives the pre-cache instruction "precache--card_id2". After parsing the instruction, it searches for the training and inference expansion card numbered 2 in the system device list, confirms that it is working properly and is well-connected to the system, and determines it as the target training and inference expansion card. The pre-cache instruction stipulates that data is prefetched according to the data usage frequency. In the storage granules of the target training and inference expansion card, the model parameter data with a higher usage frequency is located through indexing. These data are screened to remove the parts irrelevant to the current image recognition task, and the prefetch data is determined. The transmission rate of the PCIe bus is configured to 32 GB / s, and the data block size is 1 MB. The DMA transfer is started, and the prefetched model parameter data is transferred from the target training and inference expansion card to the training and inference unit through the PCIe bus.

[0124] In this embodiment, at the beginning of training, the model parameters are reasonably processed and allocated, and subsequently, according to the pre-cache instruction, the required expansion card data is loaded into the video memory of the training and inference unit, providing the necessary data support for the calculation of the training and inference unit. Combining with the subsequent training and inference data processing steps, it can effectively optimize the use of video memory for large model training and inference, reduce the demand for high-capacity video memory, and improve the overall training and inference efficiency.

[0125] Based on the above embodiments, Embodiment 5 of the present application proposes a control method for a training and inference expansion card. Before step S10, it further includes:

[0126] Step D1, determining whether the training and inference data is used for the next frame of calculation;

[0127] Step D2, if it is not used for the next frame of calculation, execute the step of determining the video memory occupancy value of the training and inference data.

[0128] In this embodiment, the next frame of calculation: during the training or inference process of the model, the calculation steps are carried out in chronological order. The next frame of calculation refers to the next calculation step immediately following the current calculation step. Whether the training and inference data is used for the next frame of calculation determines whether the data needs to be immediately retained in the video memory of the training and inference unit to meet the fast access requirement.

[0129] As an alternative implementation, the computational graph of the large model is extremely complex, containing thousands or even tens of thousands of nodes and edges. Advanced graph analysis algorithms are used to deeply analyze the computational graph to clarify the dependency relationships and usage locations of each training and inference data in the graph. For example, in large language models based on the Transformer architecture, the output data of certain intermediate layers may only be used in specific multi-head attention mechanisms or feed-forward network modules. By analyzing the computational graph, it can be accurately determined whether this data will be used in the next frame of calculation. During the training and inference processes of the large model, track the computational progress and status in real time and accurately. Utilize timestamps and event logs in the distributed system to record the generation time and usage time of each training and inference data, and combine with the current computational steps to accurately determine whether this data will be used in the next step of calculation. For example, in the distributed training of the large model, the intermediate gradient data generated by some nodes may only be used for local updates in the current iteration step and will not be used in the next frame of calculation. When it is determined that the training and inference data is not used in the next frame of calculation, perform the following operations to determine its video memory occupancy value: The data types in the large model are diverse and complex. In addition to common integers, floating-point numbers, and boolean values, there may also be custom data types. Special data type recognition algorithms are adopted to accurately identify the data types of the training and inference data. For example, for the quantized data in the large model, accurately identify its quantization bits and data format to determine the storage space occupied by each element. The data in the large model usually has extremely high dimensions. For example, the dimensions of a tensor may reach dozens or even hundreds. Efficient dimension calculation algorithms are used to accurately calculate the sizes of each dimension of the multi-dimensional data. By multiplying the number of bytes occupied by the data type by the total number of elements of the data (i.e., the product of the sizes of each dimension), the video memory occupancy value of the training and inference data is obtained. For example, for a 16-bit floating-point tensor with a shape of (1024, 1024, 1024, 64), the total number of elements is 1024×1024×1024×64, and each element occupies 2 bytes. Therefore, the video memory occupancy value of this tensor is 1024×1024×1024×64×2 bytes. When the video memory occupancy value of the training and inference data is greater than the video memory threshold, perform the following operations to store the training and inference data into the training and inference expansion card: Add a unique and detailed identifier to the training and inference data. The identifier contains rich information such as data type, generation time, the model layer it belongs to, and the data source node. At the same time, construct an efficient data index structure to facilitate subsequent rapid management and search in the training and inference expansion card. For example, use data structures such as hash tables or B-trees to store the mapping relationship between the data identifier and the storage location. Transfer the training and inference data from the video memory of the training and inference unit to the storage medium of the training and inference expansion card through a high-speed and low-latency interface (such as a PCIe5.0 or higher-level interface). Adopt advanced data compression algorithms (such as the Zstd compression algorithm) to efficiently compress the training and inference data to significantly reduce the data transmission volume and transmission time. At the same time, utilize multi-channel parallel transmission technology to further improve the transmission efficiency.Due to the huge amount of training and inference data for large models, the training and inference expansion cards may adopt a distributed storage architecture. When storing the training and inference data, a distributed file system (such as Ceph, GlusterFS, etc.) is used for management, and the data is dispersed and stored on multiple storage nodes. At the same time, the storage location and relevant meta-information (such as identification, size, checksum, etc.) of each data block are recorded to ensure the reliability and accessibility of the data.

[0130] Through the above embodiments, it is determined whether the training and inference data is used for the next frame calculation, and the unused data is stored in the training and inference expansion card according to the video memory occupancy value and the video memory threshold, effectively reducing the high-capacity requirement of the large model training for the video memory and improving the training efficiency and resource utilization rate.

[0131] Based on any of the above embodiments, Embodiment Six of the present application proposes a control method for a training and inference expansion card. The training and inference data includes at least one of model gradients, optimizer parameters, and checkpoints. The step of storing the training and inference data to the target training and inference expansion card through a direct connection bus between the training and inference unit and the target training and inference expansion card includes at least one of the following: unloading the model gradients to the training and inference expansion card; updating the optimizer state based on the optimizer parameters.

[0132] In this embodiment, the training and inference data includes at least one of model gradients, optimizer parameters, and checkpoints. Model gradients: In the training of large models, model gradients are the partial derivatives of the objective function with respect to the model parameters. It represents the rate of change of the objective function at the current parameter values, reflects in which direction the model parameters should be adjusted to minimize the objective function value, and is an important basis for the model to update parameters. Optimizer parameters: The optimizer is used to adjust the model parameters to minimize the loss function. Optimizer parameters are the parameters used by the optimizer during the update process, such as the learning rate, momentum, etc. Different optimizers have different parameter settings, and these parameters will affect the training speed and effect of the model. Checkpoint: It is a snapshot of the model during the training process, including the state of the model at a specific moment, such as the parameter values of the model, the state of the optimizer, etc. Saving checkpoints can facilitate resuming training after the training is interrupted and can also be used for model evaluation and comparison. Unloading: The process of moving data from the video memory of the training and inference unit to the training and inference expansion card, aiming to release the space of the video memory of the training and inference unit to cope with the situation of insufficient video memory. Optimizer state: The internal state of the optimizer during the training process. For example, when using the Stochastic Gradient Descent (SGD) optimizer, the accumulated value of momentum is part of the optimizer state. The optimizer state is continuously updated as the training progresses to guide the update of the model parameters.

[0133] As an alternative implementation, when the training and inference data output by the training and inference unit contains model gradients and the video memory occupancy value of the calculated model gradients is greater than the video memory threshold associated with the training and inference unit, the following operations are performed: Data tagging: Add a unique identifier to the model gradient data, which includes information such as data type (model gradient), generation time, and the model layer it belongs to, to facilitate subsequent management and search in the training and inference expansion card. Data transmission: Transfer the model gradient data from the video memory of the training and inference unit to the storage medium of the training and inference expansion card through a high-speed interface (such as a PCIe interface). During the transmission process, use a data compression algorithm (such as Huffman coding) to compress the model gradient data to reduce the data transmission volume and transmission time. Storage management: Allocate storage space for the model gradient data in the training and inference expansion card, and record its storage location and relevant metadata (such as identifier, size, etc.). The storage management can be carried out in the form of a file system or a database to facilitate quick location and access to the model gradient data. As an alternative implementation, when the training and inference data output by the training and inference unit contains optimizer parameters, the following operations are performed: Parameter parsing: Parse the optimizer parameters to extract the values of each parameter, such as learning rate, momentum, etc. Different optimizers have different parameter formats, and need to be parsed according to the specific optimizer type. State update: Update the internal state of the optimizer according to the parsed optimizer parameters. For example, when using a stochastic gradient descent optimizer with momentum, update the cumulative value of momentum according to the new momentum parameter. The update process needs to follow the update rules of the optimizer to ensure the correctness of the optimizer state. State saving: Save the updated optimizer state to the video memory of the training and inference unit or the training and inference expansion card. If the video memory space is sufficient, the optimizer state can be saved to the video memory for quick access; if the video memory space is tight, the optimizer state is saved to the training and inference expansion card and its storage location is recorded.

[0134] Optionally, after the step of unloading the model gradient to the training and inference expansion card, it includes: if a gradient update instruction from the training and inference unit is received, obtain the target model gradient corresponding to the gradient update instruction; send the target model gradient to the training and inference unit.

[0135] In this embodiment, gradient update instruction: An instruction sent by the training and inference unit to request the acquisition of a specific model gradient for gradient update operations. This instruction contains relevant identification information of the required model gradient, such as generation time, the model layer it belongs to, etc., to accurately locate the target model gradient. Target model gradient: The model gradient data that is found and determined from the training and inference expansion card according to the gradient update instruction and is required by the training and inference unit during the gradient update process. Cumulative gradient: In some deep learning training strategies, in order to simulate a larger batch size, the model gradients are gradually accumulated in multiple training steps, and then the accumulated gradients are used to update the model parameters after reaching a certain number of steps.

[0136] As an alternative implementation, the training and inference expansion card monitors the instructions from the training and inference unit in real time. When receiving a gradient update instruction, it parses the instruction and extracts the identification information of the target model gradient contained therein, such as the generation time, the model layer to which it belongs, etc. According to the parsed identification information, it searches in the storage management system of the training and inference expansion card. When the training and inference expansion card stores the model gradient data, it will record the relevant meta-information of each data (such as identification, storage location, etc.), and through these meta-information, the storage location of the target model gradient can be quickly located. After finding the target model gradient data, it verifies it to ensure the integrity and correctness of the data. Verification can be performed through methods such as checksum and hash value. If the data verification fails, corresponding error handling measures need to be taken, such as re-obtaining the data or reporting error information. The target model gradient data is read from the storage medium of the training and inference expansion card. If the data is compressed during storage, such as using Huffman coding, decompression operation needs to be performed after reading the data. The read and decompressed target model gradient data is sent to the video memory of the training and inference unit through a high-speed interface.

[0137] Optionally, the training and inference data includes model parameters. The step of storing the training and inference data to the target training and inference expansion card through the direct connection bus between the training and inference unit and the target training and inference expansion card includes: if the training and inference data further includes a checkpoint, associatively storing the checkpoint and the model parameters to the training and inference expansion card.

[0138] As an alternative implementation, the checkpoint and the model parameters are stored in the training and inference expansion card in a mutually associated manner, so that when needed in the future, the corresponding model parameters can be quickly located according to the checkpoint, or the corresponding checkpoint information can be found according to the model parameters.

[0139] Through the above embodiments, the target model gradient is obtained from the training and inference expansion card according to the gradient update instruction and sent to the training and inference unit to support the training strategy of accumulating gradients, while effectively reducing the high-capacity demand for video memory in large model training. The model parameters that do not need to participate in the calculation are offloaded to the training and inference expansion card, and the checkpoint is associatively stored with the model parameters, thereby effectively reducing the high-capacity demand for video memory in large model training and facilitating the recovery of training and the evaluation of the model.

[0140] Based on the above embodiments, Embodiment 7 of the present application proposes a control method for a training and inference expansion card, which further includes: if receiving inference data to be processed, obtaining the remaining video memory of each training and inference expansion card; if the remaining video memory is not lower than the running threshold, calling the remaining video memory to execute an inference process based on the inference data to output an inference result.

[0141] In this embodiment, the data to be inferred: In the scenario of real-time training (inference while training), it refers to the data that needs to be input into the large model for inference calculation. For example, in natural language processing tasks, the data to be inferred may be a piece of text; in image recognition tasks, it may be an image. The remaining video memory of the training and inference expansion card: It refers to the size of the unused storage space in the current video memory of the training and inference expansion card. Since the training and inference expansion card is used to store part of the training and inference data to relieve the video memory pressure of the training and inference unit, its remaining video memory reflects the available resources currently available for executing the inference process. The running threshold: It is a minimum available space standard set for the video memory of the training and inference expansion card. When the remaining video memory is not lower than this threshold, it indicates that the training and inference expansion card has sufficient resources to execute the inference process; otherwise, the inference may not be able to proceed normally. The inference process: The large model uses the learned model parameters to calculate based on the input data to be inferred, thereby obtaining the inference result. In the real-time training scenario, the inference process runs in parallel with the training process.

[0142] As an alternative implementation, the training and inference unit and the training and inference expansion card interact through a specific communication protocol (such as the PCIe interface communication protocol). The training and inference unit sends a request instruction to the training and inference expansion card to obtain the remaining video memory. After receiving the request instruction, the training and inference expansion card queries the usage of the video memory through its internal hardware management module. This module can directly access the status register of the video memory to obtain information on the used video memory space and the total video memory space. Based on the queried used video memory space and the total video memory space, the remaining video memory of the training and inference expansion card is calculated. For example, if the total video memory of the training and inference expansion card is 1024GB and 20GB has been used, then the remaining video memory is 1004GB. When the remaining video memory of the training and inference expansion card is not lower than the running threshold, the data to be inferred is loaded from the input device (such as a network interface, storage device, etc.) into the video memory of the training and inference expansion card. During the loading process, it may be necessary to preprocess the data, such as normalization, cropping, encoding, etc., to make it meet the input requirements of the large model. The relevant parameters of the large model are obtained from the storage of the training and inference expansion card and loaded into the video memory. Due to the huge number of parameters of the large model, techniques such as parameter partitioning and asynchronous loading may be adopted to improve the loading efficiency. The training and inference expansion card calls its computing core (such as a GPU core, a dedicated inference chip, etc.) to perform inference calculations based on the loaded data to be inferred and the model parameters. During the calculation process, parallel computing techniques are used to accelerate the inference speed, such as multi-threaded computing, SIMD instruction sets, etc. After the inference calculation is completed, the obtained inference result is output from the video memory of the training and inference expansion card. The result can be returned to the requester (such as a client, an application program, etc.), or further processed and analyzed.

[0143] Through the above embodiments, in the scenario of large models with real-time training (training while inferring), it is determined whether to execute the inference process based on the remaining video memory of the training and inference expansion card, giving full play to the advantage of the large capacity of the training and inference expansion card, achieving efficient inference of large models, while reducing the dependence on a large number of graphics cards, reducing costs and improving resource utilization.

[0144] The present application provides a training and inference expansion card, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the control method of the training and inference expansion card in Embodiment 1 above.

[0145] Next, refer to Figure 4 , which shows a schematic structural diagram of a training and inference expansion card suitable for implementing the embodiments of the present application. Figure 4 The training and inference expansion card shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0146] As Figure 4 shown, the training and inference expansion card may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which may execute various appropriate actions and processes according to a program stored in a read-only memory (ROM, Read Only Memory) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM, Random Access Memory) 1004. In the random access memory 1004, various programs and data required for the operation of the training and inference expansion card are also stored. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD, Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the training and inference expansion card to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows a training and inference expansion card with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems may be alternatively implemented or had.

[0147] Refer to Figure 5 , Figure 5This is another example of the training and inference expansion card in this embodiment. The training and inference expansion card includes a processing core and memory chips. The Linux system is burned into the processing core, and the memory chips are used to expand the video memory. After the training and inference expansion card is connected to the terminal, it directly interacts with the training and inference unit of the terminal by running the Linux system to train and locally deploy large models.

[0148] Specifically, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through a communication device, or installed from the storage device 1003, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the methods of the embodiments disclosed in the present application are executed.

[0149] The training and inference expansion card provided by the present application adopts the control method of the training and inference expansion card in the above embodiment, and can solve the technical problem that the local training and deployment of large models cannot be carried out due to the hardware limitations of the graphics card. Compared with the prior art, the beneficial effects of the training and inference expansion card provided by the present application are the same as those of the training and inference expansion card provided in the above embodiment, and other technical features in this training and inference expansion card are the same as those disclosed in the method of the previous embodiment, which will not be elaborated here.

[0150] It should be understood that the various parts disclosed in the present application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.

[0151] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0152] The present application provides a computer-readable storage medium with computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the control method of the training and inference expansion card in the above embodiment.

[0153] The computer-readable storage medium provided by the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in combination with an instruction execution system, device, or component. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, radio frequency (RF), etc., or any suitable combination of the above.

[0154] The above computer-readable storage medium may be included in the training and inference expansion card; or it may exist separately without being assembled into the training and inference expansion card.

[0155] The above computer-readable storage medium carries one or more programs. When the one or more programs are executed by the training and inference expansion card, the training and inference expansion card is caused to: if it detects that the training and inference unit outputs training and inference data, determine the video memory occupancy value of the training and inference data;

[0156] if the video memory occupancy value is greater than the video memory threshold associated with the training and inference unit, determine the target training and inference expansion card;

[0157] store the training and inference data to the target training and inference expansion card through the direct connection bus between the training and inference unit and the target training and inference expansion card.

[0158] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN, Local Area Network) or a wide area network (WAN, Wide Area Network), or it can be connected to an external computer (for example, by connecting through the Internet using an Internet service provider).

[0159] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of the code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutively represented blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0160] The modules described in the embodiments of this application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation on the unit itself in some cases.

[0161] The readable storage medium provided by this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., computer programs) for performing the control method of the above-mentioned training and inference expansion card, and can solve the technical problem that the local training and deployment of large models cannot be carried out due to the hardware limitations of the graphics card. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by this application are the same as those of the control method of the training and inference expansion card provided in the above embodiments, and will not be elaborated here.

[0162] An embodiment of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the steps of the control method of the training and inference expansion card as described above.

[0163] The computer program product provided by the present application can solve the technical problem that the hardware limitation of the graphics card causes the inability to locally train and deploy large models. Compared with the prior art, the beneficial effects of the computer program product provided by the embodiment of the present application are the same as those of the control method of the training and inference expansion card provided by the above embodiment, and will not be elaborated here.

[0164] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent scope of the present application.

Claims

1. A control method for a training and pushing expansion card, characterized in that: The control method of the training and pushing expansion card comprises: If it is detected that the training-pushing unit outputs the training-pushing data, determining the video memory occupancy value of the training-pushing data; If the video memory occupancy value is greater than the video memory threshold associated with the training and pushing unit, determining a target training and pushing expansion card; The training and pushing data is stored in the target training and pushing expansion card through a direct bus between the training and pushing unit and the target training and pushing expansion card.

2. The control method of the training expansion card according to claim 1, characterized in that: Before the step of determining the video memory occupancy value of the training data if it is detected that the training data unit outputs the training data, the method includes: If it is detected that at least one training-propulsion expansion card is connected, the driver configuration of the training-propulsion expansion card is loaded; Based on the driver configuration, the model of the training and pushing unit, and the hardware resources of the terminal to which the training and pushing expansion card is connected, a large model training and pushing environment is configured in the terminal.

3. The control method of the training expansion card as claimed in claim 2, characterized in that: After the step of configuring the large model training and pushing environment on the terminal, the method further includes: In response to a model training instruction, determining a model identifier corresponding to the model training instruction; A data scheduling strategy corresponding to the model identifier is determined according to the large model training environment.

4. The control method of the training expansion card as claimed in claim 3, characterized in that: After the step of determining the resource scheduling strategy according to the model identifier and the configured large model training and pushing environment, the method further includes: Get model parameters; Based on the data scheduling strategy, the model parameters are divided into video memory data and expansion card data; The expansion card data is stored in the training-pushing expansion card, and the video memory data is sent to the training-pushing unit.

5. The control method of the training expansion card as claimed in claim 4, characterized in that: Before the step of dividing the model parameters into video memory data and expansion card data based on the data scheduling strategy, the method includes: In response to a video memory configuration instruction, determining a video memory quota of a training expansion card corresponding to the video memory configuration instruction; The data scheduling strategy is updated according to the video memory quota.

6. The control method of the training expansion card as claimed in claim 1, characterized in that: The training and push expansion card includes a processing core and storage particles, the processing core is burned with a Linux system, and the control method of the training and push expansion card also includes: In response to a main card configuration instruction, determining a first training expansion card corresponding to the main card configuration instruction; The Linux system of the first training push expansion card is run, and the Linux system of the second training push expansion card is shielded, wherein the first training push expansion card and the second training push expansion card form an expansion card array.

7. The control method of the training expansion card as claimed in claim 1, characterized in that: The control method of the training and pushing expansion card also includes: In response to the pre-caching instruction, determining a target training expansion card according to the pre-caching instruction; Determining pre-fetched data in the target training and propulsion expansion card; The pre-fetched data is transmitted to the training and pushing unit through a direct bus.

8. The control method of the training expansion card as claimed in claim 1, characterized in that: Before the step of determining the video memory occupancy value of the training and pushing data, the method further includes: Determining whether the training and pushing data is used for next frame calculation; If it is not used for the next frame calculation, the step of determining the video memory occupancy value of the training and pushing data is performed.

9. A training expansion card, characterized in that: The training-push expansion card comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the control method of the training-push expansion card according to any one of claims 1 to 8.

10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the steps of the control method of the training and push expansion card according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Accelerator, memory management method for accelerator and data processing system

    CN106959893A

  • Resource dynamic adjustment method and device, equipment and storage medium

    CN110502340A

  • Heterogeneous computing system and model training method and device thereof, medium and program product

    CN118396073A

  • Searching method, searching device, computer equipment and storage medium

    CN118838917A

  • Distributed recommendation method and system based on near data processing

    CN119127727A