Terminal large model parameter collaborative loading method, scheduling method and system
By co-loading large model parameters on edge terminals, combining hard disk and remote memory loading methods, the problem of too long loading time for model parameters is solved, achieving faster loading speed and better user experience.
Patent Information
- Application Number
- CN202510607163.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-13
AI Technical Summary
When existing edge terminals perform large-scale model inference, the model parameters loading time is too long, resulting in poor user experience.
By co-loading model parameters in inference devices, using a combination of hard disk loading and remote memory loading, model parameters are divided into cached areas, hard disk loading areas and remote memory loading areas, and loading them in parallel through multithreading to improve efficiency.
It greatly improves the loading speed of model parameters on edge terminals, provides a better user experience, and does not rely on specific hardware, and has good scalability.
Smart Images

Figure CN120144490A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of large model parameter processing, and in particular to a method and system for collaborative loading and scheduling of large model parameters on a terminal. Background Art
[0002] A large language model refers to a deep learning model trained using a large amount of text data, which can generate natural language text or understand the meaning of language text. After years of development, it currently provides powerful application capabilities in the fields of natural language processing, text generation, machine translation, intelligent customer service, etc., greatly improving the level of automation and intelligence. With the continuous improvement of computer hardware performance and the emergence of stronger requirements for model understanding and expression capabilities, the scale of large model parameters is constantly expanding. Currently, with the continuous expansion of the parameter scale, local services can no longer meet the needs of intelligent applications. Most intelligent applications based on Transformer rely heavily on cloud services. As the last stop for large models to land from cloud services, edge terminals have the advantages of saving bandwidth, protecting user privacy and security, providing better offline support, lower costs, and good scalability. Therefore, performing large model inference on edge terminals will gradually become a mainstream trend. However, edge terminals often only have hardware devices with limited performance. When reading data in edge terminals, the slow hard disk reading speed in edge terminals will cause the model parameter loading time to be too long when performing large model inference on edge terminals, greatly reducing the user experience. It can be seen that the existing edge terminal parameter loading methods have problems such as too long loading time and poor user experience when performing large model inference. Summary of the Invention
[0003] The present invention provides a method and system for collaborative loading and scheduling of large model parameters on a terminal to solve the problems of too long loading time and poor user experience existing in the existing edge terminal parameter loading methods when performing large model inference.
[0004] To achieve the above object, the present invention is implemented through the following technical solutions: In a first aspect, the present invention provides a method for collaborative loading of large model parameters on a terminal, including: S110. Determine the file size according to the model file to be loaded, and based on the file size, allocate a memory area in the inference device and align the start address of the memory area byte by byte; S120. Obtain the caching situation of the model to be loaded in the memory, and divide the model to be loaded into a cached area and an uncached area according to the caching situation; S130. Obtain the hard disk loading speed of the inference device and the network situation of the remote memory device, and divide the uncached area into a hard disk loading area and a remote memory loading area; S140. Load the model parameters of the cached area from the cache of the inference device into the corresponding positions in the memory area, load the model parameters of the hard disk loading area from the hard disk of the inference device into the corresponding positions in the memory area, and load the model parameters of the remote memory loading area from the remote memory device into the corresponding positions in the memory area.
[0005] Optionally, in S110, creating a memory area in the inference device based on the file size includes: Construct a memory area size determination model, and determine the size of the memory area to be confirmed through the memory area size determination model. The model satisfies the following relational expression: ; In the formula, represents the size of the memory area to be confirmed, represents the file size of the model to be loaded, represents the alignment byte number of the tensors used by the model to be loaded, represents the memory byte number occupied by the untyped pointer in the programming language; Use the malloc function to allocate a memory area in the local memory of the inference device with the same size as the memory area to be confirmed; The byte alignment of the starting address of the memory area includes: Construct a byte alignment address determination model, and determine the byte-aligned address through the byte alignment address determination model. The model satisfies the following relational expression: ; In the formula, represents the byte-aligned address, represents the starting address of the memory area, represents the alignment byte number of the tensors used by the model to be loaded, represents the memory byte number occupied by the untyped pointer in the programming language, represents the bitwise AND operation, represents the bitwise NOT operation; Load the model parameters starting from the byte-aligned address to complete the byte alignment.
[0006] Optionally, in S120, obtaining the caching situation of the model to be loaded in the memory and dividing the model to be loaded into a cached area and an uncached area according to the caching situation includes: Obtain the caching situation of each page of the model to be loaded in the memory, mark the pages with the caching situation of cached as cached, and mark the pages with the caching situation of uncached as uncached; Logically combine the pages marked as cached into a cached area, and logically combine all the pages marked as uncached into an uncached area.
[0007] Optionally, in S130, obtain the hard disk loading speed of the inference device and the network conditions of the remote memory device, and divide the uncached area into a hard disk loading area and a remote memory loading area, including: Use a test function to obtain the local hard disk loading speed of the inference device and the bandwidth of the network connecting the remote memory device to the inference device, and calculate the ratio of the local hard disk loading speed to the network bandwidth; Calculate the total memory size required for the uncached area, and divide the total memory size required for the uncached area into a hard disk loading area and a remote memory loading area based on the ratio of the local hard disk loading speed to the network bandwidth.
[0008] Optionally, in S140, load the model parameters of the cached area from the cache of the inference device to the corresponding positions in the memory area, including: Open a first thread to load the model parameters of the cached area from the cache of the inference device to the corresponding positions in the memory area; Load the model parameters of the hard disk loading area from the hard disk of the inference device to the corresponding positions in the memory area, including: Open a second thread to load the model parameters of the hard disk loading area from the hard disk of the inference device to the corresponding positions in the memory area; Load the model parameters of the remote memory loading area from the remote memory device to the corresponding positions in the memory area, including: Open a third thread to load the model parameters of the remote memory loading area from the remote memory device to the corresponding positions in the memory area; After the model parameters of the first thread, the second thread, and the third thread are loaded, complete the parameter loading of the model file to be loaded.
[0009] In a second aspect, an embodiment of the present application provides a terminal large model parameter scheduling method for dynamically allocating the parameters of a model file to be loaded based on the terminal large model parameter collaborative loading method as described in the first aspect, including: S210. Obtain the memory usage rate, maximum memory, and network bandwidth information of all remote memory devices connected to the inference device, and calculate the parameter allocation priority of each remote memory device according to the memory usage rate, the maximum memory, and the network bandwidth information; S220. Construct a priority-based parameter allocation algorithm according to the priority assigned to the parameters, and use the priority-based parameter allocation algorithm to initially allocate the parameters of the model file to be loaded on the inference device to the memory of the remote memory device in terms of pages; S230. Obtain the memory utilization rate of the remote memory device and the timing change information of the connection status with the inference device in real time, and dynamically re-allocate the parameters of the model file to be loaded according to the timing change information by using the priority-based parameter allocation algorithm.
[0010] Optionally, in S210, calculating the parameter allocation priority of each remote memory device according to the memory utilization rate, the maximum memory, and the network bandwidth information includes: Calculating the parameter allocation priority of each remote memory device by using a parameter allocation priority calculation formula, and the calculation formula satisfies the following relational expression: ; In the formula, represents the parameter allocation priority, represents the connection situation between the remote memory device and the inference device, means there is a connection, means there is no connection, represents the maximum memory of the remote memory device, represents the network bandwidth between the remote memory device and the inference device, represents the memory utilization rate of the remote memory device; The priority-based parameter allocation algorithm includes: Dividing the parameters to be allocated into pages, and calculating the number of pages occupied by the parameters to be allocated to obtain the number of pages occupied by the parameters; Construct a priority-based parameter allocation algorithm based on the number of pages occupied by the parameters and the parameter allocation priority, and the algorithm satisfies the following relational expression: ; In the formula, represents the number of parameter pages to be allocated on each remote memory device, represents the number of pages occupied by the parameters, represents the total priority of all connected remote memory devices.
[0011] Optionally, in S230, dynamically re-allocating the parameters of the model file to be loaded according to the timing change information by using the priority-based parameter allocation algorithm includes: When it is found that the memory utilization rate of a certain remote memory device exceeds the memory utilization rate threshold, re-allocate some of the parameters on this device to other remote memory devices by using the priority-based parameter allocation method; When it is found that a certain remote memory device loses connection, all parameters on this device are reallocated to other remote memory devices using a priority-based parameter allocation method.
[0012] In a third aspect, an embodiment of the present application provides a terminal large model parameter collaborative loading system for implementing the terminal large model parameter collaborative loading method described in the first aspect, including: A system initialization module: used to initialize the inference device and the remote memory device, enabling the inference device to load data from the memory of the remote memory device, and at the same time, initially allocate the parameters of the large model file to the memory of these devices; A remote memory device management module: used to manage each remote memory device, and record in real time information such as the connection status, network bandwidth, memory usage rate, and maximum memory of each remote memory device with the inference device; A parameter management module: used to manage parameters, and record in real time the distribution of different parts of the large model parameters on different remote memory devices; A parameter scheduling decision module: used to make parameter scheduling strategy decisions and actual parameter scheduling in real time according to the remote memory device information of the remote memory device management module and the distribution of different parts of the large model parameters on different remote memory devices; A parameter loading module: used to provide the function of loading parameters from the remote memory for the program, and automatically determine the loading method according to the distribution of the large model parameters on different devices.
[0013] In a fourth aspect, an embodiment of the present application provides a terminal large model parameter scheduling system, including at least one control processor and a memory for communicating with the control processor; The memory stores instructions executable by at least one control processor, and the instructions are executed by at least one control processor, so that at least one control processor can execute the terminal large model parameter scheduling method described in the second aspect.
[0014] Beneficial effects: The terminal large model parameter collaborative loading method provided by the present invention can make more full use of the overall IO capabilities of the terminal device. On the basis of the conventional hard disk loading method, it adds the method of loading from the remote memory, and can provide a fixed multiple of acceleration effect for model files of any size (for example, when the hard disk loading speed is 80MB / s and the network bandwidth is 110MB / s, an acceleration ratio of 2.375 times can be provided). The larger the model file, the more obvious the acceleration effect. In addition, this method does not depend on any specific additional hardware, has good scalability, and can be applied to any other software or hardware acceleration method to further provide parameter loading acceleration effect.
[0015] Meanwhile, the terminal large model parameter scheduling method provided by the present invention can make more full use of the idle memory of surrounding terminal devices in the edge terminal scenario. Since the memory utilization rate, maximum memory, and network bandwidth factors of the remote memory device are used to determine the parameter allocation priority, this method can ensure optimization goals such as minimizing the consumption of remote memory and minimizing the mapping and switching of model parameters with the remote memory device, which can reduce the waste of remote memory and transmission bandwidth. In addition, the load balancing of the remote memory device is also considered to prevent excessive memory pressure on some remote memory devices. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 is a schematic flowchart of the terminal large model parameter collaborative loading method provided by the present invention; Figure 2 is a data flow diagram of the terminal large model parameter collaborative loading method provided by the present invention; Figure 3 is a schematic flowchart of the terminal large model parameter scheduling method provided by the present invention; Figure 4 is a structural diagram of the terminal large model parameter collaborative loading system provided by the present invention; Figure 5 is a schematic structural diagram of the terminal large model parameter scheduling system provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] The technical solutions of the present invention will be described clearly and completely below. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0018] Unless otherwise defined, the technical terms or scientific terms used in the present invention shall have the ordinary meaning understood by those of ordinary skill in the art to which the present invention belongs. The "first", "second" and similar terms used in the present invention do not indicate any order, quantity or importance, but are only used to distinguish different components. Similarly, terms such as "a" or "one" do not indicate a quantity limitation, but indicate the existence of at least one. "Connection" or "connected" and similar terms are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right", etc. are only used to indicate relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship also changes accordingly.
[0019] Please refer to Figure 1-2 , as Figure 1As shown in the figure, an embodiment of the present application provides a method for collaborative loading of large model parameters based on remote memory. The method at least includes the following steps: S110. Obtain the file size according to the model file to be loaded by the inference device, manually allocate a memory area of a certain size based on the file size, and align the start address of the area byte by byte; According to the model file path, use the fopen function to open the model file in the "rb" mode, use the fseek function to move the file read / write pointer to the end of the file, use the ftell function to obtain the offset of the file read / write pointer relative to the start of the file at this moment. This offset is equal to the file size, and the file size is denoted as filesize. Finally, use the rewind function to make the file read / write pointer return to the start of the file; Use the getpagesize function to obtain the page size supported by the system, denoted as pagesize, and the alignment byte number used by the large model is denoted as alignment. According to the formula obtain the size of the memory area to be allocated, memorysize. In the formula, sizeof(void *) refers to the untyped pointer in the programming language, that is, the number of bytes of the memory space occupied by the pointer of the void type in the C++ language. Use the malloc function to allocate a memory area of size memorysize in the memory, and save the start address of this memory area to the raw pointer; Use the formula Align raw to the modelbuffer address according to the alignment requirement. Subsequently, use modelbuffer as the start address for loading large model parameters, and at the same time assign the value of the raw pointer to the modelbuffer[-1] pointer to facilitate subsequent release of the entire memory area allocated by malloc.
[0020] S120. Obtain the caching situation of the model file in the memory, and divide the model file into a cached area and an uncached area; According to the model file path, use the open function to open the file in read-only mode to obtain the file descriptor of the model file, and use the mmap function to map the model file to the memory address space according to the file descriptor; Use the mincore function to obtain the caching status of each page of the model file in memory. Create two vector containers, the cached page vector and the uncached page vector. Traverse the caching status of all pages in sequence and perform different operations according to whether the page is cached. If the page is already cached, store the start address and end address of some parameters of the page in the cached page vector. If the page is not cached, store the start address and end address of some parameters of the page in the uncached vector.
[0021] S130. Obtain the hard disk loading speed of the inference device and the network condition of the remote memory device, and further divide the uncached area into a hard disk loading area and a remote memory loading area; Use a self-written test function to obtain the hard disk loading speed of the inference device and the network speed of the remote memory device (i.e., the speed of loading parameters from the remote memory). Divide the local hard disk loading speed by the remote memory loading speed to obtain a ratio, denoted as M; Calculate the number of pages in the uncached page vector, and divide the pages into two parts according to the ratio of M. The first part of the pages is further divided into the hard disk loading page vector, and the second part of the pages is divided into the remote memory loading page vector.
[0022] S140. Load the model parameters in the cached area from the cache of the inference device to the corresponding positions in the manually allocated memory area, load the model parameters in the hard disk loading area from the hard disk of the inference device to the corresponding positions in the manually allocated memory area, and load the model parameters in the remote memory loading area from the remote memory to the corresponding positions in the manually allocated memory area.
[0023] Create a thread and use the fread function to load the parameters pointed to by the cached page vector into the corresponding positions in the manually allocated memory area. At this time, the fread function uses the file descriptor of the model file in the cache of the inference device to read the parameters; Create a thread and use the fread function to load the parameters pointed to by the hard disk loading page vector into the corresponding positions in the manually allocated memory area. At this time, the fread function uses the file descriptor of the model file in the hard disk of the inference device to read the parameters; Create a thread and use the fread function to load the parameters pointed to by the cached page vector into the corresponding positions in the manually allocated memory area. At this time, the fread function uses the file descriptor of the model file in the remote memory to read the parameters; Use the pthread_join function for each thread to wait for them to finish execution. After all threads have finished execution, the model parameter loading process is completed, and the starting address of the model file in memory is returned to the program for subsequent inference processes.
[0024] It can be understood that the method for collaborative loading of large model parameters based on remote memory can provide a stable parameter loading acceleration effect under stable network conditions. After experimental tests, when the hard disk loading speed of the inference device is 80 MB / s and the network bandwidth between the inference device and the remote memory device is 110 MB / s, using this parameter collaborative loading method can increase the original parameter loading speed of 80 MB / s to 190 MB / s, providing a stable acceleration ratio of 2.375 times. When using a model file of 3.6 GB, the parameter loading time can be reduced from 44 seconds to 18 seconds, greatly reducing the waiting time for users to perform large model inferences.
[0025] Please refer to Figure 3 , this application embodiment provides a large model parameter scheduling method based on remote memory, which is used to support the method for collaborative loading of large model parameters based on remote memory for terminals, including: S210. Obtain the memory usage rate, maximum memory, and network bandwidth information of all remote memory devices connected to the inference device, and calculate the parameter allocation priorities of each of the remote memory devices according to the memory usage rate, maximum memory, and network bandwidth information of the remote memory devices; Use the ping command to obtain the connection status connect between the inference device and each remote memory device. When the connect value is 1, it means there is a network connection between the inference device and this device; when it is 0, it means there is no network connection between the inference device and this device. Use the free command and calculate the memory usage rate mem_usage and maximum memory max_mem of each remote memory. Start the iperf server on all remote memory devices, and use the iperf client program on the inference device to measure the network bandwidth net_speed of each remote memory device respectively. Then, according to the formula Obtain the priority P of each remote memory device i .
[0026] S220. According to the parameter allocation priorities of each of the remote memory devices, use a priority-based parameter allocation algorithm to initially allocate the parameters of the model file on the inference device to the memory of the remote memory devices in page granularity; Record the number of pages occupied by the model file as N. Then, according to the formula Calculate that the number of parameters to be allocated on each remote memory is S i pages, and then round S i up and down to make Si The values are all integers, and the total number is equal to the total number of pages of the model file; Traverse each remote memory device in turn. According to the number of pages S to be allocated for them i to determine the range of model file parameters that should be stored in their memory, and load the corresponding part of the model parameters into their memory in advance through the network.
[0027] S230. Obtain the memory usage rate of the remote memory device and the timing change information of the connection status with the inference device in real time. According to the timing change information, use the parameter allocation algorithm based on priority to dynamically re-allocate the large model parameters.
[0028] Use the free command every 1 second and calculate the memory usage rate mem_usage and the maximum memory max_mem of each remote memory, and use the ping command to judge the connection status between the inference device and other remote memory devices; If it is found that the mem_usage of a certain remote memory device exceeds the specified threshold (such as 80%), try to re-allocate the model parameters on the memory of this device to other remote memory devices. According to the formula Calculate the priority P of other remote memory devices i , and judge whether releasing the memory of a part of the parameters of this device can reduce the memory usage rate to below the specified threshold (such as 60%). If so, then re-allocate this part of the parameters to the corresponding remote memory device according to the parameter allocation priority of other remote memory devices, and release the memory occupied by this part of the parameters on this remote memory device. If not, then re-allocate all the parameters to the corresponding remote memory device according to the parameter allocation priority of other remote memory devices, and release the memory occupied by this part of the parameters on this remote memory device.
[0029] If it is found that a certain remote memory device loses connection, then re-allocate the model parameters on this device to the corresponding remote memory device according to the parameter allocation priority of other remote memory devices.
[0030] Repeat the above process.
[0031] Please refer to Figure 4 , the embodiment of the present application provides a terminal large model parameter collaborative loading system, including: System initialization module 310: used to initialize the inference device and the remote memory device, so that the inference device can load data from the memory of the remote memory device, and at the same time initially allocate the parameters of the large model file to the memory of these devices; The specific initialization process includes: Create a memory file system using tmpfs on each remote memory device and mount a part of the memory space as a virtual memory disk; On each host, use the NFS technology to share the directory mounted by tmpfs over the network; The inference device accesses the data on the remote memory by mounting the directories shared by each connectable remote memory device through the NFS technology; Allocate and load the model parameters to each remote memory device according to the aforementioned large model parameter scheduling method based on remote memory.
[0032] Remote memory device management module 320: used to manage each remote memory device, and record in real time information such as the connection status, network bandwidth, memory usage rate, and maximum memory of each remote memory device and the inference device; Parameter management module 330: used to manage parameters, and record in real time the distribution of different parts of the large model parameters on different remote memory devices; Parameter scheduling decision module 340: used to make parameter scheduling strategy decisions and actual parameter scheduling in real time according to the remote memory device information of the remote memory device management module and the distribution of different parts of the large model parameters on different remote memory devices; Parameter loading module 350: used to provide the program with the function of loading parameters from the remote memory, and automatically determine the loading method according to the distribution of the large model parameters on different devices.
[0033] This module will accept the program's request to load parameters from the remote memory, and according to the information of the parameter management module, determine the distribution of different parameters on different remote memory devices, and convert a request to load parameters from the remote memory into requests to read different parameters from different remote memory devices.
[0034] For example, when the program issues a request to load the parameters on pages 1 to 10 from the remote memory, after receiving the request, this module will access the parameter distribution information of the parameter management module. Assuming that at this time, it is found that the parameters on pages 1 to 3 are stored on remote memory device 1, the parameters on pages 4 to 8 are stored on remote memory device 2, and the parameters on pages 9 to 10 are stored on remote memory device 3. At this time, this module will convert the request to load the parameters on pages 1 to 10 from the remote memory into three requests: loading pages 1 to 3 from remote memory device 1, loading pages 4 to 8 from remote memory device 2, and loading pages 9 to 10 from remote memory device 3, and use different threads to perform the loading.
[0035] Please refer to Figure 5 , this application embodiment provides a terminal large model parameter scheduling system, including: At least one memory; At least one control processor; At least one program; The program is stored in the memory, and the control processor executes at least one program to implement the method for collaborative loading of terminal large model parameters based on remote memory and the method for scheduling large model parameters based on remote memory described in the embodiments of the present disclosure.
[0036] The electronic device can be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA), an in-vehicle computer, etc.
[0037] The following details the electronic device according to the embodiments of the present application.
[0038] The control processor 410 can be implemented by using a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present disclosure; The memory 420, where one memory is to be implemented in the form of a non-volatile memory (NVMe), and the other memory is to be implemented in the form of a random access memory (RAM). The memory 420 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 420 and are called by the control processor 410 to execute the method for collaborative loading of terminal large model parameters based on remote memory and the method for scheduling large model parameters based on remote memory of the embodiments of the present disclosure.
[0039] The input / output interface 430 is used to implement information input and output; The communication interface 440 is used to implement communication interaction between this device and other devices, and can implement communication through a wired manner (such as USB, network cable, etc.) or through a wireless manner (such as mobile network, WIFI, Bluetooth, etc.); The bus 450 transmits information between the various components of the device (such as the control processor 410, the memory 420, the input / output interface 430, and the communication interface 440); Among them, the control processor 410, the memory 420, the input / output interface 430, and the communication interface 440 are communicatively connected to each other inside the device through the bus 450.
[0040] Embodiments of the present disclosure also provide a storage medium, which is a computer-readable storage medium storing computer-executable instructions for causing a computer to execute the above-mentioned method for collaborative loading of large model parameters based on remote memory and the method for scheduling large model parameters based on remote memory.
[0041] As a non-transitory computer-readable storage medium, a memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include memories remotely located relative to the control processor, and these remote memories can be connected to the control processor through a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0042] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative efforts. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field of the present invention through logical analysis, reasoning, or limited experiments based on the concept of the present invention on the basis of the prior art should fall within the protection scope determined by the claims.
Claims
1. A method for collaboratively loading terminal large model parameters, characterized in that: include: S110, determining a file size according to the model file to be loaded, and based on the file size, opening a memory area in the inference device, and byte-aligning a start address of the memory area; S120, obtaining a cache status of the model to be loaded in the memory, and dividing the model to be loaded into a cached area and a non-cached area according to the cache status; S130, obtaining the hard disk loading speed of the inference device and the network status of the far memory device, and dividing the uncached area into a hard disk loading area and a far memory loading area; S140, load the model parameters of the cached area from the cache of the inference device to the corresponding position of the memory area, load the model parameters of the hard disk loading area from the hard disk of the inference device to the corresponding position of the memory area, and load the model parameters of the far memory loading area from the remote memory device to the corresponding position of the memory area.
2. The terminal large model parameter collaborative loading method according to claim 1 is characterized in that: In S110, the memory area is allocated in the inference device based on the file size, including: A memory area size determination model is constructed to determine the size of the memory area to be confirmed. The memory area size determination model satisfies the following relationship: ; In the formula, Indicates the size of the memory area to be confirmed. Indicates the file size of the model to be loaded, Indicates the number of aligned bytes used by the model to be loaded. Indicates the number of bytes of memory occupied by untyped pointers in programming languages; Use the malloc function to allocate a memory area in the local memory of the inference device that is the same size as the memory area to be confirmed; The byte alignment of the starting address of the memory area includes: A byte alignment address determination model is constructed, and the address after byte alignment is determined by the byte alignment address determination model. The byte alignment address determination model satisfies the following relationship: ; In the formula, Indicates the byte-aligned address. Indicates the first address of the memory area. Indicates the number of aligned bytes used by the model to be loaded. Indicates the number of bytes of memory occupied by untyped pointers in programming languages. Represents a bitwise AND operation, Represents bitwise inverse operation; Load the model parameters starting from the byte-aligned address to complete the byte alignment.
3. The terminal large model parameter collaborative loading method according to claim 1 is characterized in that: In S120, the cache status of the model to be loaded in the memory is obtained, and the model to be loaded is divided into a cached area and a non-cached area according to the cache status, including: Get the cache status of each page of the model to be loaded in memory, and mark the pages that are cached as cached, and mark the pages that are not cached as not cached; The pages marked as cached are logically combined into a cached area, and all pages marked as uncached are logically combined into an uncached area.
4. The terminal large model parameter collaborative loading method according to claim 1 is characterized in that: In S130, obtaining the hard disk loading speed of the inference device and the network status of the far memory device, and dividing the uncached area into a hard disk loading area and a far memory loading area, includes: Use the test function to obtain the local hard disk loading speed of the inference device and the bandwidth of the network connecting the remote memory device and the inference device, and calculate the ratio of the local hard disk loading speed to the network bandwidth; The total memory size to be occupied by the uncached area is calculated, and the total memory size to be occupied by the uncached area is divided into a hard disk loading area and a remote memory loading area based on the ratio of the local hard disk loading speed to the network bandwidth.
5. The terminal large model parameter collaborative loading method according to claim 1 is characterized in that: In S140, loading the model parameters of the cached area from the cache of the inference device to the corresponding position of the memory area includes: Opening a first thread to load the model parameters of the cached area from the cache of the inference device to the corresponding position of the memory area; Loading the model parameters of the hard disk loading area from the hard disk of the inference device to the corresponding position of the memory area includes: Opening a second thread to load the model parameters of the hard disk loading area from the hard disk of the inference device to the corresponding position of the memory area; Loading the model parameters of the remote memory loading area from the remote memory device to the corresponding position of the memory area includes: Opening a third thread to load the model parameters of the remote memory loading area from the remote memory device to the corresponding position of the memory area; When the model parameters of the first thread, the second thread and the third thread are loaded, the loading of the parameters of the model file to be loaded is completed.
6. A terminal large model parameter scheduling method, used for dynamically allocating parameters of a model file to be loaded based on the terminal large model parameter collaborative loading method according to any one of claims 1 to 5, characterized in that: include: S210, obtaining memory usage, maximum memory, and network bandwidth information of all far memory devices connected to the inference device, and calculating the parameter allocation priority of each far memory device according to the memory usage, the maximum memory, and the network bandwidth information; S220, constructing a priority-based parameter allocation algorithm according to the parameter allocation priority, and using the priority-based parameter allocation algorithm to preliminarily allocate the parameters of the model file to be loaded on the inference device to the memory of the far memory device at a page granularity; S230, acquiring in real time the memory usage rate of the far memory device and the time series change information of the connection status with the inference device, and dynamically reallocating the parameters of the model file to be loaded using a priority-based parameter allocation algorithm according to the time series change information.
7. The terminal large model parameter scheduling method according to claim 6, characterized in that: In S210, the parameter allocation priority of each far memory device is calculated according to the memory usage rate, the maximum memory and the network bandwidth information, including: The parameter allocation priority calculation formula is used to calculate the parameter allocation priority of each far memory device, and the calculation formula satisfies the following relationship: ; In the formula, Indicates the priority of parameter assignment, Indicates the connection status between the far memory device and the inference device. For the connection to exist, If there is no connection, Indicates the maximum memory of the far memory device. represents the network bandwidth between the far memory device and the inference device, Indicates the memory usage of the far memory device; The priority-based parameter allocation algorithm comprises: The parameters to be allocated are divided into pages, and the number of pages of the parameters to be allocated is calculated to obtain the number of pages occupied by the parameters; A priority-based parameter allocation algorithm is constructed based on the number of pages occupied by the parameters and the parameter allocation priority, and the algorithm satisfies the following relationship: ; In the formula, Indicates the number of parameter pages that need to be allocated on each far memory device. Indicates the number of pages occupied by the parameter. Represents the sum of the priorities of all connected far memory devices.
8. The terminal large model parameter scheduling method according to claim 6, characterized in that: In S230, dynamically reallocating the parameters of the model file to be loaded using a priority-based parameter allocation algorithm according to the time series change information includes: When it is found that the memory usage of a far memory device exceeds the memory usage threshold, some parameters on the device are reallocated to other far memory devices using a priority-based parameter allocation method; When it is found that a far memory device loses connection, all parameters on the device are reallocated to other far memory devices using a priority-based parameter allocation method.
9. A terminal large model parameter collaborative loading system, used to implement the terminal large model parameter collaborative loading method according to any one of claims 1 to 5, characterized in that: include: System initialization module: used to initialize the inference device and the far memory device, so that the inference device can load data from the memory of the far memory device, and at the same time, preliminarily allocate the parameters of the large model file to the memory of these devices; Far memory device management module: used to manage each far memory device, and record the connection status, network bandwidth, memory usage, maximum memory, and other information of each far memory device and inference device in real time; Parameter management module: used to manage parameters and record the distribution of different parts of large model parameters on different far memory devices in real time; Parameter scheduling decision module: used to make parameter scheduling strategy decisions and actual parameter scheduling in real time based on the far memory device information of the far memory device management module and the distribution of different parts of the large model parameters on different far memory devices; Parameter loading module: used to provide the program with the function of loading parameters from remote memory, and automatically determine the loading method based on the distribution of large model parameters on different devices.
10. A terminal large model parameter scheduling system, characterized in that: comprising at least one control processor and a memory for communicatively connecting with the control processor; The memory stores instructions that can be executed by at least one control processor, and the instructions are executed by at least one control processor to enable the at least one control processor to execute the terminal large model parameter scheduling method as described in claims 6-8.
Citation Information
Patent Citations
Distributed training and reasoning method, system and device based on artificial intelligence, and readable storage medium
CN114035937A
File loading method, computing device and storage medium
CN114706828A
Model execution method and device, storage medium and equipment
CN119025259A
Model file loading method and system, computer equipment and storage medium
CN119620958A
External memory as an extension to local primary memory
US20210240616A1