Model loading method and device, storage medium and electronic equipment

By using model header information to determine tensor data information in the AI inference system, loading it into shared memory in parallel and mapping it to the target process, the problem of low model loading efficiency in traditional AI inference systems is solved, and efficient model loading and resource utilization are achieved.

CN120276791AInactive Publication Date: 2025-07-08JINAN INSPUR DATA TECH CO LTD

Patent Information

Application Number
CN202510768696.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-07-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The problem of low model loading efficiency in traditional AI inference systems.

Method used

The data information of the tensor data is determined based on the model header information of the model to be loaded, and loaded it in parallel to the shared memory of the target node, and then map the tensor data to the target process for loading, and parallel loading is achieved using thread pool and memory mapping technology.

Benefits of technology

It improves model loading efficiency, reduces cold start time and resource consumption, and is suitable for computing-intensive and data-intensive AI model loading scenarios, improving the performance and resource utilization of AI inference services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120276791A_ABST
    Figure CN120276791A_ABST
Patent Text Reader

Abstract

The invention discloses a model loading method and device, a storage medium and electronic equipment, and relates to the technical field of computers. Data information of tensor data of a to-be-loaded model is determined according to model head information of the to-be-loaded model; the tensor data can be loaded in parallel into the shared memory of the target node of the model through the data information, and then the tensor data stored in the shared memory is mapped into the target process for loading the to-be-loaded model, so that the to-be-loaded model can be loaded through the target process. As the tensor data can be loaded into the shared memory in parallel and then mapped into the target process capable of loading the to-be-loaded model, large-scale parallel loading can be realized, the model can be directly loaded from the shared memory, and the model data does not need to be read from a disk, so that the loading efficiency of the model is improved. The technical problem of low model loading efficiency can be solved, and the technical effect of improving the model loading efficiency is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly to a method, apparatus, storage medium, and electronic device for loading a model. Background Art

[0002] In the related art, a traditional AI inference system usually loads model data from a disk, resulting in low model loading efficiency.

[0003] Regarding the above problems existing in the related art, no effective solution has been proposed yet. Summary of the Invention

[0004] This application provides a method, apparatus, storage medium, and electronic device for loading a model, so as to at least solve the problem of low model loading efficiency existing in the related art.

[0005] This application provides a method for loading a model, including: determining data information of tensor data of a model to be loaded based on model header information of the model to be loaded; parallelly loading the tensor data into a shared memory of a target node based on the data information, where the target node is a node for running the model to be loaded; mapping the tensor data stored in the shared memory to a target process for loading the model to be loaded; and loading the model to be loaded through the target process.

[0006] In an exemplary embodiment, determining data information of tensor data of a model to be loaded based on model header information of the model to be loaded includes: parsing the model header information to determine an offset and a size of the tensor data; and determining the offset and the size as the data information of the tensor data.

[0007] In an exemplary embodiment, parallelly loading the tensor data into a shared memory of a target node based on the data information includes: parallelly reading the tensor data from a hard disk in the target node according to the offset and the size of the tensor data included in the data information; and copying the tensor data into the shared memory.

[0008] In an exemplary embodiment, parallelly reading the tensor data from a hard disk in the target node according to the offset and the size of the tensor data included in the data information includes: determining a reading thread for reading the tensor data; and controlling the reading thread to parallelly read the tensor data from the hard disk in the target node according to the offset and the size of the tensor data.

[0009] In an exemplary embodiment, determining a reading thread for reading the tensor data includes: when a processor of the target node includes one core, determining the reading thread from the one core; and when the processor includes multiple cores, determining the reading thread from the multiple cores.

[0010] In an exemplary embodiment, before mapping the tensor data stored in the shared memory to the target process for loading the model to be loaded, the method further includes: creating a function instance when there is no instance for loading the model to be loaded in the target node; creating a target process in the function instance.

[0011] In an exemplary embodiment, before determining the data information of the tensor data of the model to be loaded based on the model header information of the model to be loaded, the method further includes: determining a first model required to respond to the inference request when receiving an inference request; determining the first model as the model to be loaded.

[0012] In an exemplary embodiment, before determining the first model required to respond to the inference request, the method further includes: creating a target application for providing an inference service in the target platform; receiving an inference request sent through the target application.

[0013] In an exemplary embodiment, before determining the data information of the tensor data of the model to be loaded based on the model header information of the model to be loaded, the method further includes: determining the historical access data of the inference models deployed in the target platform; determining the hot models included in the inference models based on the historical access data, where the access volume of the hot models is greater than that of other models, and the other models are the models other than the hot models included in the inference models; determining the hot models as the models to be loaded.

[0014] In an exemplary embodiment, the method further includes: determining the non-hot models included in the shared memory, where the non-hot models are the models with an access volume less than a preset threshold; clearing the non-hot models from the shared memory.

[0015] In an exemplary embodiment, the method further includes: determining the historical performance data of the system running the inference models deployed in the target platform, and the historical access data of the inference models; determining the peak load period of the system and the target access volume during the peak load period based on the historical performance data and the historical access data; determining the first number of instances for loading the models in the target node; performing a capacity adjustment operation on the target node based on the target access volume and the first number.

[0016] In an exemplary embodiment, performing a capacity adjustment operation on the target node based on the target access volume and the first number includes: determining a second number corresponding to the target access volume, where the second number is used to indicate the number of instances required when the access volume of the model is the target access volume; performing an expansion operation on the target node when the second number is greater than the first number; performing a contraction operation on the target node when the second number is less than the first number.

[0017] In an exemplary embodiment, the operation of expanding the target node includes: creating an instance of an expansion function in the target node within a predetermined time period before the peak load period; initializing and loading the instance of the expansion function so that the instance of the expansion function is in a warm-up state.

[0018] In an exemplary embodiment, after performing a capacity adjustment operation on the target node based on the target access volume and the first quantity, the method further includes: when the capacity adjustment operation is an expansion operation and the peak load period is reached, sending the access traffic to the instance of the expansion function in the warm-up state, so as to load the model indicated by the access traffic through the instance of the expansion function.

[0019] In an exemplary embodiment, after sending the access traffic to the instance of the expansion function in the warm-up state, the method further includes: determining a third quantity of instances required for the access traffic; when the third quantity is greater than the quantity of instances for loading the model indicated by the access traffic, creating a supplementary function instance in the target node; after creating the supplementary function instance is completed, sending a part of the traffic included in the access traffic to the supplementary function instance, so as to load the model indicated by the part of the traffic through the supplementary function instance.

[0020] In an exemplary embodiment, mapping the tensor data stored in the shared memory to the target process for loading the model to be loaded includes: establishing a direct mapping relationship between the tensor data in the shared memory and the virtual memory address space of the target process; allowing the target process to directly access the tensor data based on the direct mapping relationship through a memory pointer.

[0021] This application also provides a model loading device, including: a determination module, configured to determine data information of tensor data of the model to be loaded based on the model header information of the model to be loaded; a first loading module, configured to parallelly load the tensor data into the shared memory of the target node based on the data information, where the target node is the node running the model to be loaded; a mapping module, configured to map the tensor data stored in the shared memory to the target process for loading the model to be loaded; a second loading module, configured to load the model to be loaded through the target process.

[0022] This application also provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any of the above model loading methods when executing the computer program.

[0023] This application also provides a computer-readable storage medium, where a computer program is stored in the computer-readable storage medium, and where the computer program implements the steps of any of the above model loading methods when being executed by a processor.

[0024] The present application also provides a computer program product, including a computer program which, when executed by a processor, implements the steps of any one of the above-mentioned model loading methods.

[0025] Through the present application, the data information of the tensor data of the model to be loaded can be determined according to the model header information of the model to be loaded first. Through the data information, the tensor data can be loaded into the shared memory of the target node of the model with a running cost in parallel, and then the tensor data stored in the shared memory is mapped to the target process for loading the model to be loaded. Furthermore, the model to be loaded can be loaded through the target process. Since the tensor data can be loaded into the shared memory in parallel and then the tensor data is mapped to the target process that can load the model to be loaded, that is, large-scale parallel loading can be achieved, which can replace the traditional millisecond-level disk I / O (Input / Output), and the model can be directly loaded from the shared memory without reading the model data from the disk. Therefore, the technical problem of low model loading efficiency can be solved, and the technical effect of improving the model loading efficiency can be achieved. Description of the Drawings

[0026] In order to more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0027] Figure 1 It is a hardware structure block diagram of a computer terminal on which the method embodiment of the present invention runs;

[0028] Figure 2 It is a flowchart of the model loading method according to the embodiment of the present invention;

[0029] Figure 3 It is a structure block diagram of the model loading device according to a specific embodiment of the present invention;

[0030] Figure 4 It is a structure block diagram of the data caching device according to the embodiment of the present invention. Detailed Embodiments

[0031] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0032] It should be noted that in the description of this application, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0033] To enable those skilled in the art of this technology to better understand the solution of this application, the following further detailed description of this application will be given in conjunction with the accompanying drawings and specific embodiments.

[0034] In combination with the specific application environment architecture or specific hardware architecture on which the execution of the loading method of the combined model depends, the specific application environment architecture or specific hardware architecture will be described herein.

[0035] The method embodiments provided in the embodiments of this application can be executed on a mobile terminal, a computer terminal or a similar computing device. Taking the example of running on a computer terminal, Figure 1 is the hardware structure block diagram of the computer terminal on which the method embodiment of the present invention runs. As Figure 1 shown, the computer terminal may include one or more ( Figure 1 only one is shown in the figure) processors 102 (the processors 102 may include, but are not limited to, processing devices such as a microprocessor MCU (Microcontroller Unit) or a field-programmable gate array FPGA (Field-Programmable Gate Array)) and a memory 104 for storing data. Among them, the above computer terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above computer terminal. For example, the computer terminal may further include more or fewer components than Figure 1 shown in the figure, or have a different configuration from Figure 1 shown in the figure.

[0036] The memory 104 can be used to store computer programs, such as software programs and modules of application software, such as the computer program corresponding to the model loading method in the embodiments of the present invention. The processor 102 executes various functional applications and data processing by running the computer programs stored in the memory 104, that is, the above-mentioned method is implemented. The memory 104 may include high-speed random access memory and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 may further include memories remotely disposed relative to the processor 102, and these remote memories can be connected to the computer terminal through a network. Examples of the above-mentioned network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and combinations thereof.

[0037] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by a communication provider of a computer terminal. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (Radio Frequency, abbreviated as RF) module, which is used to communicate with the Internet wirelessly.

[0038] Embodiments of the present application provide a model loading method, and the method is described in detail in combination with the execution flow of the model loading method.

[0039] The following is an explanation of the professional terms that appear in the present application:

[0040] Serverless awareness: Serverless computing is a new paradigm of cloud computing. Traditional server-aware computing is resource-centric, and developers need to manage cloud resources based on application requirements. The new serverless computing is function-centric, and developers only need to focus on writing cloud functions to implement application logic, and the cloud resource management is completely responsible for the cloud computing system software. Serverless computing simplifies the process of developers writing and deploying cloud applications, can automatically expand and contract according to application requirements, and provides a fine-grained billing model, thus saving the operating cost of cloud applications.

[0041] Serverless instance: Refers to a computing unit running in a serverless computing architecture. Usually, it is a computing unit dynamically allocated by the cloud platform when a certain event is triggered and used to execute specific tasks.

[0042] Memory mapping: The process of mapping the disk sectors of a file to the virtual memory space of a process, that is, mapping a file to the virtual space of a process to achieve a one-to-one correspondence between the disk address of the file and a section of virtual addresses in the virtual space of the process. After achieving such a mapping relationship, the process can read and write this section of memory in the form of pointers, and the system will automatically write back the dirty pages to the corresponding file disk, that is, the operation on the file is completed without having to call system call functions such as read and write.

[0043] In this embodiment, a method for loading a model is provided. Figure 2 It is a flowchart of the method for loading a model according to an embodiment of the present invention, as Figure 2 shown, and this process includes the following steps:

[0044] Step S202, determine the data information of the tensor data of the model to be loaded based on the model header information of the model to be loaded;

[0045] Step S204, parallelly load the tensor data into the shared memory of the target node based on the data information, where the target node is the node that runs the model to be loaded;

[0046] Step S206, map the tensor data stored in the shared memory to the target process for loading the model to be loaded;

[0047] Step S208, load the model to be loaded through the target process.

[0048] In the above embodiment, in the case of receiving the instruction for issuing the model to be loaded from the user, the initial part of the model to be loaded can be read first to obtain the complete Header JSON (JavaScript Object Notation) data (i.e., the above-mentioned model header information). Since the amount of Header data is small and contains the key information required for subsequent operations, reading the initial part of the model to be loaded and obtaining the complete Header JSON are usually executed synchronously. Among them, Header JSON can be understood as a kind of JSON-format metadata at the head of a file or data packet, which can be used to describe the basic information, structure, and format of the file or data packet. Parsing the model header information Header JSON can obtain all the data information of the tensor data, and through the data information, the tensor information can be stored in the shared memory area based on RAM (such as / dev / shm) in a parallel and efficient manner. The loaded content can preferably ensure that the model data that is continuously or periodically accessed frequently resides in the memory to support fast access, multi-instance sharing, and provide an efficient data basis for high-concurrency inference.

[0049] In the above embodiments, the user can also create a function instance, which can map the tensor data in the shared memory to the virtual memory address space of the target process for loading the model to be loaded when starting up, and then can load the model to be loaded in the target process.

[0050] In the above embodiments, the model loading method can be used in a serverless computing architecture, such as a Serverless instance platform.

[0051] Through the present application, the data information of the tensor data of the model to be loaded can be determined according to the model header information of the model to be loaded first. Through the data information, the tensor data can be loaded in parallel into the shared memory of the target node of the running cost in the model, and then the tensor data stored in the shared memory is mapped to the target process for loading the model to be loaded, and then the model to be loaded can be loaded through the target process. Since the tensor data can be loaded in parallel into the shared memory and then the tensor data is mapped to the target process that can load the model to be loaded, that is, large-scale parallel loading can be realized, which can replace the traditional millisecond-level disk I / O (Input / Output), and the model can be directly loaded from the shared memory without reading the model data from the disk. Therefore, the technical problem of low model loading efficiency can be solved, and the technical effect of improving the model loading efficiency can be achieved.

[0052] In an exemplary embodiment, determining the data information of the tensor data of the model to be loaded based on the model header information of the model to be loaded includes: parsing the model header information to determine the offset and size of the tensor data; determining the offset and size as the data information of the tensor data.

[0053] In the above embodiments, the model header information is usually stored in a certain structured data format (such as JSON), and contains various metadata information of the model, including the offset and size of the tensor data. Therefore, parsing the model header information can obtain the "map" of all tensor data, that is, the specific position (offset and size) of each tensor in the file. The offset and size of the tensor data can be determined as the data information of the tensor data, which can be used for subsequent asynchronous parallelism.

[0054] In an exemplary embodiment, loading the tensor data in parallel into the shared memory of the target node based on the data information includes: reading the tensor data in parallel from the hard disk in the target node according to the offset and size of the tensor data included in the data information; copying the tensor data into the shared memory.

[0055] In the above embodiments, according to the parsed model header information, the reading task of each tensor data can be decomposed into independent task units. That is, "reading the data block of each tensor data" can be defined as an independent and parallelizable task unit. Each task unit includes the offset and size of the tensor data to be read. Then, the read tensor data can be copied into the shared memory. After copying the read tensor data into the shared memory, multiple Serverless instances can access the same model data, reducing the repeated loading of the model and saving memory resources at the same time.

[0056] In an exemplary embodiment, parallelly reading tensor data from the hard disk in the target node according to the offset and size of the tensor data included in the data information includes: determining a reading thread for reading the tensor data; controlling the reading thread to parallelly read the tensor data from the hard disk in the target node according to the offset and size of the tensor data.

[0057] In the above embodiments, since the Header can provide the exact byte range of each tensor data block (e.g., [offset, offset + size)), the reading operations of different tensor data are independent of each other and there is no interference. Therefore, an efficient loading can be achieved by creating a thread pool. That is, for each (or a batch of) tensor data listed in the Header, a "reading task" (i.e., the above reading thread) can be generated and submitted to the thread pool. Each thread can parallelly read a specified length of byte data from a specified position on the hard disk in the target node according to the specified offset and size of the file. In a multi-core CPU or high-concurrency environment, this mechanism can make full use of computing resources and greatly improve the loading efficiency.

[0058] In an exemplary embodiment, determining a reading thread for reading tensor data includes: in the case where the processor in the target node includes one core, determining the reading thread from one core; in the case where the processor includes multiple cores, determining the reading thread from multiple cores.

[0059] In the above embodiments, the number of cores in the processor of the target node can be determined. If the processor contains only one core, the read thread for reading tensor data can be determined from this core. Since all read tasks will be executed by a single or multiple threads on this one core, the thread pool technology can be used to pre-create a certain number of threads (e.g., equal to or less than the number of cores), and assign the read tasks to these threads to maximize the processing capacity of a single core. If the processor contains multiple cores, the read threads for reading tensor data can be determined from these cores. The thread allocation can take into account the parallel processing ability of multiple cores to achieve higher read efficiency. That is, one or more read threads can be created for each core, and the read tasks can be distributed to different cores according to the number of cores and the computational intensity of the tasks. For example, thread affinity settings can be used to bind specific threads to specific cores, which can reduce cache switching between CPUs and improve the read speed. Based on the number of cores of the processor and the load condition of the current task, the system can dynamically adjust the number of read threads. On a multi-core processor, the number of threads can be equal to or less than the number of cores to ensure that there are sufficient resources to execute the read operation without causing excessive scheduling. If the read tasks are relatively few compared to the number of processor cores, the system can reduce the number of threads in the thread pool to reduce the overhead of thread context switching. Once the read threads are determined, the system will assign the read tasks to these threads. The thread scheduler is responsible for allocating CPU time slices to ensure that each read thread has the opportunity to execute its task. On a multi-core processor, different read threads can execute simultaneously, thus achieving true parallel processing. The system will continuously monitor the status of each read thread to ensure that they are running normally, and update the status in the shared memory after completing the read task to prepare for subsequent data copying and access. Through the number of cores of the processor and the task characteristics, the system can effectively determine and allocate read threads to achieve parallel reading of tensor data, significantly improving the efficiency of AI model loading, reducing the cold start time and resource consumption. This method is particularly applicable to the scenarios of loading computationally intensive and data-intensive AI models, and can make full use of the high performance of modern multi-core processors and the high concurrency ability of high-speed storage devices.

[0060] In an exemplary embodiment, before mapping the tensor data stored in the shared memory to the target process for loading the model to be loaded, the method further includes: in the case where there is no instance for loading the model to be loaded in the target node, creating a function instance; creating a target process in the function instance.

[0061] In the above embodiments, it is possible to detect whether there is a Serverless instance in the target node that matches the model to be loaded. If the user fails to find an available instance, the scaling-up process can be triggered, that is, a target process can be constructed in the created function instance. After creating the function instance, a running environment can be prepared for the new instance, which can include installing necessary software libraries, configuring running parameters, setting environment variables, etc., to ensure that the instance can execute smoothly when loading the model. By creating new function instances and target processes, continuous and reliable inference services can be provided in scenarios with high concurrency or frequent model switching. The above strategy for dynamic instance creation and process initialization makes full use of the elasticity and flexibility of the Serverless architecture. At the same time, combined with shared memory and memory mapping technologies, the model loading time can be greatly reduced, and the resource utilization rate and overall service performance can be improved.

[0062] In an exemplary embodiment, before determining the data information of the tensor data of the model to be loaded based on the model header information of the model to be loaded, the method further includes: when receiving an inference request, determining a first model required to respond to the inference request; and determining the first model as the model to be loaded.

[0063] In the above embodiments, when receiving one or more inference requests from a user, the first model required to respond to the inference request can be determined according to the model identifier in the request. Among them, the inference request can be sent from the client to the Serverless platform through an API gateway, an HTTP interface, or other communication protocols, and can include information such as a model identifier and input data, for indicating an inference task for executing a specific model. The model can be any type of AI model, such as a deep learning model, a machine learning model, etc. By effectively responding to the inference request, it can be ensured that the correct model is used for processing, and the model can be quickly loaded when needed, providing low-latency and high-efficiency inference services for users.

[0064] In an exemplary embodiment, before determining the first model required to respond to the inference request, the method further includes: creating a target application for providing inference services in the target platform; and receiving an inference request sent through the target application.

[0065] In the above embodiments, the user can create a target application for providing AI model inference services through the graphical interface of the Serverless platform (i.e., the above-mentioned target platform), send inference requests through the API gateway and generate access traffic. In a Serverless environment, from creating the application to receiving the inference request, and then to model determination and loading warm-up, a complete closed-loop can be formed, ensuring the high efficiency and low latency of the AI inference service, being able to dynamically adjust resources according to the real-time request volume, provide on-demand services, while also reducing resource waste and waiting time, and improving the user experience and system throughput.

[0066] In the above embodiments, after creating the target application, it can be deployed to the Serverless platform, which may include uploading the application code to the storage system of the platform, and at the same time, instances can be automatically allocated and configured according to the resource requirements of the application. During the deployment process, necessary model files can be pre-loaded into the high-speed storage of the target node for subsequent rapid access. After the target application is deployed, the system can receive the inference requests sent through this application. Among them, the inference requests can be sent to the application through HTTP API, message queue or other communication mechanisms. When the request arrives, it can be routed to the corresponding function instance, which may contain the model identifier and the data required for inference. After receiving the inference request, the front-end routing logic of the target application will check the model identifier in the request to determine the first model required to respond to this inference request. If the model has not been loaded on any instance, or the current instance cannot meet the concurrent requirements of the request, the system will trigger a new model loading or instance expansion process.

[0067] In an exemplary embodiment, before determining the data information of the tensor data of the model to be loaded based on the model header information of the model to be loaded, the method further includes: determining the historical access data of the inference models deployed in the target platform; determining the hot models included in the inference models based on the historical access data, where the access volume of the hot models is greater than that of other models, and other models are models other than the hot models included in the inference models; determining the hot models as the models to be loaded.

[0068] In the above embodiments, the running data (i.e., the above-mentioned historical access data) can be continuously collected from sources such as the target platform, API gateway, and function instances, such as timestamps, model identifiers, request frequencies, durations, etc. Through the above historical access data, the user access patterns and the popularity of specific models can be identified. Through the popularity of different specific models, the hot models with the highest access volume can be determined, and the hot models are determined as the models to be loaded. The collected raw data can also be cleaned, aggregated, and structured, and stored in a data storage suitable for time series analysis or pattern mining.

[0069] In the above embodiment, the hot spot model is intelligently determined and preloaded based on the historical access data, which can significantly improve the response speed and resource utilization of the model reasoning. It can not only process the current request, but also predict the future load and allocate resources in advance, thus avoiding the cold start delay in high concurrency and improving the overall Serverless reasoning service performance.

[0070] In an exemplary embodiment, the method further includes: determining a non-hotspot model included in the shared memory, wherein the non-hotspot model is a model with a visit volume less than a preset threshold; and clearing the non-hotspot model in the shared memory.

[0071] In the above embodiment, the number of accesses of each inference model in the shared memory can be continuously counted, and a preset threshold of the number of accesses can be set based on business needs and resource management strategies to distinguish between hotspot models and non-hotspot models. Among them, the preset threshold can be set based on the average access frequency of the model, the percentile of the number of accesses, the number of accesses in a specific time window, etc. By comparing the number of accesses of the model with the preset threshold, non-hotspot models with low accesses can be identified, that is, this non-hotspot model may occupy shared memory resources due to its low frequency of use, resulting in resource waste. Therefore, after identifying the non-hotspot model, in order to maintain the efficient use of shared memory, a strategy for clearing the non-hotspot model can be implemented to release the memory space occupied by it, so as to reserve more resources for possible hotspot models or high concurrent requests in the future. Among them, the system performs a clearing operation to remove the non-hotspot model from the shared memory, which may include updating the memory management data structure, releasing the relationship between the model data and the memory mapping, releasing the relevant memory pages, and updating the model status information. After clearing the non-hotspot model, the released memory space can be used to store new hotspot model data, or to expand the number of instances of the existing model to cope with higher concurrent requests. The system dynamically adjusts resource allocation strategies based on current load conditions and predicted model access patterns.

[0072] In the above embodiment, through dynamic resource optimization of shared memory, it is possible to ensure that memory space is efficiently utilized while reducing resource waste and cold start delays. It is particularly suitable for scenarios where model access patterns change dynamically, and can automatically adjust the loading strategy according to the popularity of the model, thereby improving the flexibility and response speed of AI reasoning services under the Serverless architecture. At the same time, by intelligently managing shared memory resources, the system can meet real-time reasoning requirements while maintaining high resource utilization, providing users with a better service experience.

[0073] In an exemplary embodiment, the method further includes: determining historical performance data of a system running an inference model deployed in a target platform, and historical access data of the inference model; determining a peak load period of the system and a target access volume during the peak load period based on the historical performance data and the historical access data; determining a first number of instances for loading the model in a target node; and performing a capacity adjustment operation on the target node based on the target access volume and the first number.

[0074] In the above embodiment, machine learning methods can be used to analyze the stored historical data to identify key patterns, such as: the periodicity of the request volume, the trend of the popularity change of a specific model, the characteristics of burst traffic, etc.; based on the historical data, a prediction model can be used to generate prediction results for a future period of time, including: the overall request load trend, the expected request volume of high-frequency models, potential traffic peak periods, etc. Therefore, by continuously collecting performance and request data from each target process, storing them in a database, and periodically running analysis and prediction algorithms.

[0075] In the above embodiment, based on the prediction results of the above prediction algorithm and the real-time AI inference request number, CPU usage rate or memory usage rate (i.e., the above historical performance data), the peak load period of the system and the corresponding target access volume can be predicted. Through the target access volume and the first number of instances for loading the model, the number and resource configuration of Serverless instances can be adjusted to achieve on-demand expansion and contraction. The system can intelligently predict and respond to the resource requirements during the peak load period. By dynamically adjusting the number of instances and the model preloading strategy, it can provide an efficient and fast-response inference service, while avoiding resource waste and achieving a good balance between cost and performance.

[0076] In an exemplary embodiment, performing a capacity adjustment operation on the target node based on the target access volume and the first number includes: determining a second number corresponding to the target access volume, where the second number is used to indicate the number of instances required when the access volume of the model is the target access volume; performing an expansion operation on the target node when the second number is greater than the first number; and performing a contraction operation on the target node when the second number is less than the first number.

[0077] In the above embodiments, when the number of instances required to reach the target traffic volume (i.e., the second number above) is less than the first number, it can be indicated that the existing resources are insufficient to handle the traffic volume during peak hours, and thus an expansion operation needs to be performed; if the second number is greater than the first number, it can be indicated that the current number of instances exceeds the actual demand, and a scaling-down operation can be performed to reduce resource waste. By intelligently responding to changing model access requirements and dynamically adjusting the number of instances according to the predicted target traffic volume, both the response speed and processing capacity of the service are ensured, and over-allocation and waste of resources are avoided, achieving an efficient and economical Serverless inference service.

[0078] In an exemplary embodiment, the operation of expanding the target node includes: creating an expanded function instance in the target node within a predetermined time period before the peak load period; initializing and loading the expanded function instance to make the expanded function instance in a warm-up state.

[0079] In the above embodiments, additional Serverless function instances can be created in the target node within a predetermined time period before the predicted peak load period. The predetermined time can be 30 minutes, 40 minutes, 50 minutes, etc., but is not limited thereto. The newly created expanded function instances need to be initialized, which can include loading the runtime environment, dependency libraries, configuration files, etc. Subsequently, the model loading process can be executed. Through asynchronous parallel loading technology, the tensor data of the hot model can be efficiently loaded from high-speed storage to shared memory, ensuring that the expanded instances can quickly access the model data. After the expanded function instances complete the model loading, they will enter the warm-up state. During the entire warm-up period, the real-time traffic volume and resource usage can be continuously monitored. If there are deviations between the actual situation and the prediction, the number of instances can be dynamically adjusted according to the latest data to better match the actual demand. By creating and warming up the expanded function instances in advance, the resource pressure during the peak load period can be effectively alleviated, and the stability and response speed of the AI inference service are improved. At the same time, the instances in the warm-up state do not cause resource waste during non-peak hours, achieving effective management and optimal utilization of resources.

[0080] In an exemplary embodiment, after performing the capacity adjustment operation on the target node based on the target traffic volume and the first number, the method further includes: when the capacity adjustment operation is an expansion operation and the peak load period is reached, sending the traffic to the expanded function instances in the warm-up state to load the model indicated by the traffic through the expanded function instances.

[0081] In the above embodiments, when the actual traffic reaches its peak, the pre-warmed standby instances can be directly awakened, skipping the cold start and model loading phases. The instances will immediately switch from the standby state to the active state to process inference requests. By making full use of the dynamic scaling feature of the Serverless architecture and combining pre-loading and pre-warming strategies, on-demand resource allocation and performance optimization can be achieved.

[0082] In an exemplary embodiment, after sending access traffic to the scaled-out function instances in the pre-warmed state, the method further includes: determining a third quantity of instances required for the access traffic; creating supplementary function instances in the target node when the third quantity is greater than the quantity of instances for loading the model indicated by the access traffic; after creating the supplementary function instances is completed, sending a portion of the traffic included in the access traffic to the supplementary function instances to load the model indicated by the portion of the traffic.

[0083] In the above embodiments, the current access traffic can be evaluated in real time, and based on the current traffic, model processing requirements, and instance load conditions, the third quantity of required instances can be calculated. If the third quantity is greater than the number of instances currently loading the model (including the scaled-out instances in the pre-warmed state), the creation of supplementary function instances can be triggered. These newly created instances will also be started in the target node, following the same initialization and model loading process to ensure that they can be ready to process requests in a short time. The newly created supplementary function instances also need to enter the pre-warmed state. Once the supplementary function instances are created and the initialization is completed, a portion of the access traffic can be reallocated to these new instances considering the processing capabilities and current loads of each instance. The traffic allocation will consider the processing capabilities and current loads of each instance to achieve optimal resource utilization and the shortest request response time. Among them, multiple scaled-out function instances can share the same shared memory data to reduce duplicate loading and memory redundancy.

[0084] In an exemplary embodiment, mapping the tensor data stored in the shared memory to the target process for loading the model to be loaded includes: establishing a direct mapping relationship between the tensor data in the shared memory and the virtual memory address space of the target process; allowing the target process to directly access the tensor data based on the direct mapping relationship through a memory pointer.

[0085] In the above embodiment, when the function instance is started, the memory mapping function can be called, and a direct mapping relationship between the tensor data in the shared memory and the reserved memory area of ​​the target process can be established. The process can directly access the data in the shared memory through the virtual address without the need for additional data copying or transmission processes, which greatly improves the speed and efficiency of data access. When the mapping relationship is established, the target process can directly access the tensor data through the memory pointer. When accessing the tensor data in the shared memory through memory mapping, the cache mechanism can be used to further improve the performance. For example, when the process accesses a certain memory address for the first time, the data will be loaded into the cache; subsequent accesses can be read directly from the cache, thereby avoiding frequent physical memory access and improving the access speed of the data. When the target process no longer needs to access specific tensor data, it can release the memory mapping and release this part of the virtual memory space without affecting the data in the shared memory. In addition, if multiple processes access the same data in the shared memory at the same time, it can ensure that the read and write operations of the data are synchronized to avoid data inconsistency problems.

[0086] The following is an explanation of the model loading method in conjunction with a specific embodiment:

[0087] Figure 3 is a structural block diagram of a model loading device according to a specific embodiment of the present invention, such as Figure 3 As shown, the device comprises:

[0088] Function reasoning application component 302 is deployed in the Serverless computing platform as an entry point for providing AI model reasoning services to the outside world.

[0089] Model Block Module ( Figure 3(not shown in the figure), during the loading process, the initial part of the model file (i.e., the model to be loaded mentioned above) is first read to obtain the complete Header JSON (i.e., the model header information mentioned above). This step is usually executed synchronously because the Header data volume is small and contains key information required for subsequent operations. By parsing the Header, the system can obtain the "map" of all tensors (i.e., the data information mentioned above), that is, the specific position (offset and size) of each tensor in the file. Then, "reading the data block of each tensor" is defined as an independent and parallelizable task unit. Since the Header has provided the exact byte range of each tensor data block (e.g., [offset, offset + size)), the reading operations of different tensors are independent of each other without interference. Finally, the module can achieve efficient loading by creating a thread pool. For each (or a batch of) tensors listed in the Header, the system module submits a "reading task" to the thread pool. Each task is responsible for reading a specified length of byte data from the specified offset of the file. In a multi-core CPU or high-concurrency environment, this mechanism can make full use of computing resources and greatly improve the loading efficiency.

[0090] The shared memory module 306 is responsible for efficiently storing AI model data (such as deep learning model files) in the RAM-based shared memory area. The loading content of this module can be guided by the intelligent warm-up strategy of the scaling module, and it gives priority to ensuring that the model data with continuous or periodic high-frequency access resides in memory to support fast access, multi-instance sharing, and provide an efficient data foundation for high-concurrency inference.

[0091] The data collection module 308 continuously collects operation data (such as timestamps, model identifiers, request frequencies, durations, etc.) from sources such as the Serverless platform, API gateways, and function instances, and identifies user access patterns and the popularity of specific models. The raw data collected is cleaned, aggregated, and structured, and stored in a data store suitable for time series analysis or pattern mining.

[0092] The intelligent prediction module 310 uses machine learning methods to analyze the stored historical data and identify key patterns, such as: the periodicity of request volumes, the trend of popularity changes of specific models, the characteristics of burst traffic, etc.; based on historical data, it uses a prediction model to generate prediction results for a future period of time, including: the overall request load trend, the expected request volume of high-frequency models, potential traffic peak periods, etc., and outputs the analysis results and prediction data to the scaling module as the input for its decision-making.

[0093] The scaling module 312 is responsible for dynamically adjusting the number and resource configuration of Serverless instances based on the analysis insights and prediction results provided by the intelligent prediction module, as well as the real-time AI inference request count, CPU usage, or memory usage, so as to achieve on-demand expansion and contraction. It supports the warm-up strategy, pre-starts some instances to reduce cold start latency, and ensures that the model data in the shared memory is ready to optimize performance and cost.

[0094] In the above embodiment, a Serverless computing platform is deployed and provided, which supports multi-instance parallel execution of AI inference tasks. The data collection module and the intelligent prediction module continuously collect system operation data in the background for analysis and prediction. When initializing a Serverless function or container, first synchronously read the model file header and parse the Header JSON to obtain the tensor position information (offset and size). Subsequently, the reading of each tensor data block is defined as an independent task and executed in parallel through a thread pool. Each task reads a fixed number of bytes of data from the specified offset without interfering with each other. In a multi-core CPU or high-concurrency environment, computing resources can be fully utilized to greatly improve the loading efficiency. After the function instance loads the model, it executes the inference task and returns the result to the user. When the inference request load is high, the scaling module can perform rapid expansion according to the output of the intelligent prediction module (such as predicted load peaks and hot models) and real-time monitoring data. Multiple Serverless instances share the same model data in the shared memory and quickly load the inference model to start the inference service.

[0095] In the above embodiments, the model chunking module and the memory sharing module can be deployed on each node. After receiving the model distribution instruction from the user or based on the prediction result of the prediction module, the initial part of the model file can be read first to obtain the complete Header JSON. By parsing the Header, the system can obtain the "map" of all tensors, that is, the specific position (offset and size) of each tensor in the file. Then, "reading the data block of each tensor" is defined as an independent and parallelizable task unit. Finally, the module can create a thread pool. For each (or batch) of tensors listed in the Header, the system module can submit a "reading task" to the thread pool. Each task makes full use of the multi-core computing resources of the CPU to read the specified length of byte data from the specified offset of the NVMe (Non-Volatile Memory Express) high-speed storage and copy it into the shared memory, giving priority to ensuring the availability of the hot model data for prediction; The user creates a function application for providing AI model inference services through the graphical interface of the Serverless platform; The user sends an inference request through the API gateway to generate access traffic. If no available instance is found, the scaling process is triggered to create a function instance. When the function instance starts, it directly maps the model file in the shared memory to the virtual memory address space of the process with the help of the memory mapping module, allowing the process to directly access the data through the memory pointer without any data copying. The data collection and intelligent prediction module continuously collects the performance and request data from each component, stores them in the database, and periodically runs analysis and prediction algorithms to update the hot model list and future load prediction. The intelligent prediction module, according to the predicted peak load, triggers the pre-start of the instance 30 minutes in advance. The scaling module scales the capacity on demand and binds the shared memory model data, and the instance is in a warm-up state. When the actual traffic reaches the peak, the pre-warmed standby instance is directly awakened, skipping the cold start and model loading phases, and the instance immediately switches from the standby state to the active state to process the inference request; If there are not enough pre-warmed instances, real-time scaling is triggered to supplement new instances. Multiple scaled function instances share the same shared memory data, reducing duplicate loading and memory redundancy.

[0096] In the foregoing embodiments, by asynchronously and in parallel reading a specified length of byte data from a specified offset of an NVMe high-speed storage through a thread pool and efficiently copying it into the shared memory, the cold start latency and resource consumption of Serverless functions can be significantly reduced. At the same time, multiple function instances can share the same physical memory page in the shared memory, which can avoid repeated loading of model data and reduce memory redundancy. On this basis, an intelligent warm-up mechanism is also introduced: it can identify hot models according to historical access patterns and ensure their priority and continuous residence in the shared memory; it can also analyze the request pattern to predict upcoming load peaks and automatically expand and warm up the required number of Serverless instances in advance, effectively avoiding the cold start problem. Therefore, through asynchronous parallel technology and combined with the on-demand rapid expansion and intelligent warm-up strategy technology of serverless computing, high-speed response, high resource utilization, and high scalability of the inference service are achieved, greatly saving manpower and maintenance costs and having high application value.

[0097] In the foregoing embodiments, by utilizing the 64K parallel queue feature supported by the NVMe protocol, the large model is sliced into multiple 4MB data blocks (Blocks) according to the tensor dimension, and high-speed access to the model data by Serverless instances is achieved through multi-channel concurrent loading, replacing the traditional millisecond-level disk I / O. This enables the instance to quickly load the model and perform inference. When scaling out in a high-concurrency scenario, all instances share a single model copy in the shared memory, greatly reducing memory occupancy and redundant loading. Based on the NVMe asynchronous parallel loading technology, a specified offset and length of byte data of the large model file are efficiently read from the NVMe high-speed storage and quickly copied into the shared memory to achieve rapid model loading. At the same time, combined with the on-demand expansion feature of serverless computing and the intelligent warm-up strategy, based on historical data, the model with high-frequency calls is preferentially pre-loaded into the shared memory, and the peak load is predicted, and Serverless instances are started in advance to dynamically optimize resource allocation, which can solve the cold start latency and resource redundancy problems of AI model inference in the serverless environment. Through the description of the above implementation manners, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware, but in many cases, the former is a better implementation manner.

[0098] In the foregoing embodiments, by asynchronously and parallelly preloading AI model data into the shared memory, the cold start latency of Serverless function instances can be significantly reduced, the response speed can be increased, and it is ensured that even in high-concurrency scenarios, the instances can be quickly started and begin to process inference requests. Multiple function instances share a single copy of the model data in the memory, avoiding memory and computing resource redundancy caused by repeated loading of model data, achieving efficient utilization of resources, and reducing the overall operating cost. Combining with the intelligent prediction preheating technology, the present invention can dynamically adjust the number of Serverless instances, start the instances in advance to adapt to the expected load peak, ensure the scalability of the system, and can quickly respond to changing traffic demands. The serverless computing model simplifies the programming process of developers, without the need to pay attention to the details of underlying resource management and scheduling, and focuses on writing business logic, improving the development efficiency and the maintainability of the code. By intelligently analyzing historical access patterns and predicting future loads, the present invention can precisely control the preheating and startup of instances, avoid resource waste caused by over-preheating, achieve performance maximization while optimizing cost expenditure. During peak load periods, the preheated instances can be immediately put into service, reducing the risk of service interruption caused by insufficient resources and enhancing the overall stability and reliability of the service. For large AI models, the loading method proposed by the present invention can effectively process model files of GB level, solve the common performance bottleneck problems in large-scale model deployment, and is applicable to processing complex and large-scale data sets.

[0099] An embodiment of the present application further provides a model loading device, Figure 4 which is a structural block diagram of a data caching device according to an embodiment of the present invention, as Figure 4 shown. The device includes:

[0100] A determination module 42, configured to determine data information of tensor data of a model to be loaded based on model header information of the model to be loaded;

[0101] A first loading module 44, configured to parallelly load tensor data into the shared memory of a target node based on the data information, where the target node is a node running the model to be loaded;

[0102] A mapping module 46, configured to map the tensor data stored in the shared memory to a target process for loading the model to be loaded;

[0103] A second loading module 48, configured to load the model to be loaded through the target process,

[0104] wherein, the determination module 42 corresponds to the above-mentioned model chunking module, the first loading module 203 corresponds to the above-mentioned shared memory module 306, and the mapping module 46 and the second loading module 48 correspond to the above-mentioned function inference application component 302.

[0105] In an exemplary embodiment, the determining module 42 may determine the data information of the tensor data of the model to be loaded based on the model header information of the model to be loaded in the following manner: parse the model header information to determine the offset and size of the tensor data; determine the offset and size as the data information of the tensor data.

[0106] In an exemplary embodiment, the first loading module 44 may load the tensor data into the shared memory of the target node in parallel based on the data information in the following manner: read the tensor data in parallel from the hard disk in the target node according to the offset and size of the tensor data included in the data information; copy the tensor data into the shared memory.

[0107] In an exemplary embodiment, the first loading module 44 may read the tensor data in parallel from the hard disk in the target node according to the offset and size of the tensor data included in the data information in the following manner: determine the read thread for reading the tensor data; control the read thread to read the tensor data in parallel from the hard disk in the target node according to the offset and size of the tensor data.

[0108] In an exemplary embodiment, the first loading module 44 may determine the read thread for reading the tensor data in the following manner: when the processor in the target node includes one core, determine the read thread from one core; when the processor includes multiple cores, determine the read thread from multiple cores.

[0109] In an exemplary embodiment, before mapping the tensor data stored in the shared memory to the target process for loading the model to be loaded, the apparatus: when there is no instance for loading the model to be loaded in the target node, create a function instance; create a target process in the function instance.

[0110] In an exemplary embodiment, before determining the data information of the tensor data of the model to be loaded based on the model header information of the model to be loaded, the apparatus: when receiving an inference request, determine a first model required to respond to the inference request; determine the first model as the model to be loaded.

[0111] In an exemplary embodiment, before determining the first model required to respond to the inference request, the apparatus: create a target application for providing inference services in the target platform; receive an inference request sent through the target application.

[0112] In an exemplary embodiment, the apparatus is further configured to, before determining the data information of the tensor data of the model to be loaded based on the model head information of the model to be loaded: determine the historical access data of the inference model deployed in the target platform; determine the hot models included in the inference model based on the historical access data, where the access volume of the hot models is greater than that of other models, and the other models are the models other than the hot models included in the inference model; and determine the hot models as the models to be loaded.

[0113] In an exemplary embodiment, the apparatus may also be configured to: determine the non-hot models included in the shared memory, where the non-hot models are the models with an access volume less than a preset threshold; and clear the non-hot models in the shared memory.

[0114] In an exemplary embodiment, the apparatus may also be configured to: determine the historical performance data of the system running the inference model deployed in the target platform, and the historical access data of the inference model; determine the peak load period of the system and the target access volume during the peak load period based on the historical performance data and the historical access data; determine the first number of instances in the target node for loading the model; and perform a capacity adjustment operation on the target node based on the target access volume and the first number.

[0115] In an exemplary embodiment, the apparatus may implement the capacity adjustment operation on the target node based on the target access volume and the first number in the following manner: determine the second number corresponding to the target access volume, where the second number is used to indicate the number of instances required when the access volume of the model is the target access volume; perform an expansion operation on the target node when the second number is greater than the first number; and perform a contraction operation on the target node when the second number is less than the first number.

[0116] In an exemplary embodiment, the apparatus may implement the expansion operation on the target node in the following manner: create an expansion function instance in the target node within a predetermined time period before the peak load period; initialize and load the expansion function instance so that the expansion function instance is in a warm-up state.

[0117] In an exemplary embodiment, after performing the capacity adjustment operation on the target node based on the target access volume and the first number, the apparatus is further configured to: when the capacity adjustment operation is an expansion operation and the peak load period is reached, send the access traffic to the expansion function instance in the warm-up state to load the model indicated by the access traffic through the expansion function instance.

[0118] In an exemplary embodiment, the apparatus is further configured to, after sending access traffic to an expanded function instance in a warm-up state: determine a third quantity of instances required for the access traffic; create supplementary function instances in the target node when the third quantity is greater than the quantity of instances for loading the model indicated by the access traffic; and after creating the supplementary function instances, send a portion of the traffic included in the access traffic to the supplementary function instances to load the model indicated by the portion of the traffic through the supplementary function instances.

[0119] In an exemplary embodiment, the mapping module 46 may map the tensor data stored in the shared memory to the target process for loading the model to be loaded in the following manner: establish a direct mapping relationship between the tensor data in the shared memory and the virtual memory address space of the target process; and allow the target process to directly access the tensor data through a memory pointer based on the direct mapping relationship.

[0120] For the description of the features in the corresponding embodiment of the model loading apparatus, reference may be made to the relevant description in the corresponding embodiment of the model loading method, which will not be elaborated here one by one.

[0121] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above-described embodiments of the model loading method.

[0122] An embodiment of the present application further provides a computer-readable storage medium in which a computer program is stored. The computer program is configured to execute the steps in any of the above-described embodiments of the model loading method when running.

[0123] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disc that can store a computer program.

[0124] An embodiment of the present application further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any of the above-described embodiments of the model loading method are implemented.

[0125] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above-described embodiments of the model loading method are implemented.

[0126] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0127] The above has introduced in detail a method, device, storage medium, and electronic device for loading a model provided by this application. Specific examples are used herein to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A method for loading a model, characterized in that Including: Determining data information of tensor data of the model to be loaded based on model header information of the model to be loaded; Parallelly loading the tensor data into the shared memory of a target node based on the data information, where the target node is the node running the model to be loaded; Mapping the tensor data stored in the shared memory to a target process for loading the model to be loaded; Loading the model to be loaded through the target process.

2. The model loading method according to claim 1, wherein: The determining data information of tensor data of the model to be loaded based on model header information of the model to be loaded includes: Parsing the model header information to determine the offset and size of the tensor data; Determining the offset and the size as the data information of the tensor data.

3. The model loading method according to claim 1, wherein: The parallelly loading the tensor data into the shared memory of a target node based on the data information includes: Parallelly reading the tensor data from a hard disk in the target node according to the offset and size of the tensor data included in the data information; Copying the tensor data into the shared memory.

4. The model loading method according to claim 3, wherein: The parallelly reading the tensor data from a hard disk in the target node according to the offset and size of the tensor data included in the data information includes: Determining a reading thread for reading the tensor data; Controlling the reading thread to parallelly read the tensor data from a hard disk in the target node according to the offset and size of the tensor data.

5. The model loading method according to claim 4, wherein: The determining a reading thread for reading the tensor data includes: When a processor of the target node includes one core, determining the reading thread from the one core; When the processor includes multiple cores, determining the reading thread from the multiple cores.

6. The model loading method according to claim 1, wherein: Before mapping the tensor data stored in the shared memory to a target process for loading the model to be loaded, the method further includes: When there is no instance for loading the model to be loaded in the target node, creating a function instance; Creating the target process in the function instance.

7. The model loading method according to claim 1, wherein: Before determining data information of tensor data of the model to be loaded based on model header information of the model to be loaded, the method further includes: When receiving an inference request, determining a first model required to respond to the inference request; Determining the first model as the model to be loaded.

8. The model loading method according to claim 7, wherein: Before determining a first model required to respond to the inference request, the method further includes: Creating a target application for providing an inference service in a target platform; Receive the inference request sent through the target application.

9. The method for loading a model according to claim 1, wherein before determining the data information of the tensor data of the model to be loaded based on the model header information of the model to be loaded, the method further includes: determine the historical access data of the inference model deployed in the target platform; determine the hot model included in the inference model based on the historical access data, wherein the access volume of the hot model is greater than that of other models, and the other models are models other than the hot model included in the inference model; determine the hot model as the model to be loaded.

10. The method for loading a model according to claim 1, wherein the method further includes: determine the non-hot models included in the shared memory, wherein the non-hot models are models with an access volume less than a preset threshold; clear the non-hot models in the shared memory.

11. The method for loading a model according to claim 1, wherein the method further includes: determine the historical performance data of the system running the inference model deployed in the target platform, and the historical access data of the inference model; determine the peak load period of the system and the target access volume during the peak load period based on the historical performance data and the historical access data; determine the first number of instances in the target node for loading the model; perform a capacity adjustment operation on the target node based on the target access volume and the first number.

12. The method for loading a model according to claim 11, wherein the performing a capacity adjustment operation on the target node based on the target access volume and the first number includes: determine the second number corresponding to the target access volume, wherein the second number is used to indicate the number of instances required when the access volume of the model is the target access volume; perform an expansion operation on the target node when the second number is greater than the first number; perform a contraction operation on the target node when the second number is less than the first number.

13. The method for loading a model according to claim 11, wherein the performing an expansion operation on the target node includes: create an expansion function instance in the target node within a predetermined time period before the peak load period; initialize and load the expansion function instance so that the expansion function instance is in a warm-up state.

14. The method for loading a model according to claim 11, wherein after performing a capacity adjustment operation on the target node based on the target access volume and the first number, the method further includes: when the capacity adjustment operation is an expansion operation and reaches the peak load period, send the access traffic to the expansion function instance in the warm-up state to load the model indicated by the access traffic through the expansion function instance.

15. The method for loading a model according to claim 14, wherein After sending access traffic to the scaled-out function instance in the warm-up state, the method further includes: Determining a third quantity of instances required for the access traffic; When the third quantity is greater than the quantity of instances for loading the model indicated by the access traffic, creating supplementary function instances in the target node; After creating the supplementary function instances is completed, sending a part of the traffic included in the access traffic to the supplementary function instances to load the model indicated by the part of the traffic through the supplementary function instances.

16. The model loading method according to claim 1, wherein: Mapping the tensor data stored in the shared memory to a target process for loading the model to be loaded includes: Establishing a direct mapping relationship between the tensor data in the shared memory and the virtual memory address space of the target process; Allowing the target process to directly access the tensor data based on the direct mapping relationship through a memory pointer.

17. A loading device for a model, characterized in that Includes: A determination module for determining data information of the tensor data of the model to be loaded based on the model header information of the model to be loaded; A first loading module for parallelly loading the tensor data into the shared memory of the target node based on the data information, wherein the target node is the node running the model to be loaded; A mapping module for mapping the tensor data stored in the shared memory to a target process for loading the model to be loaded; A second loading module for loading the model to be loaded through the target process.

18. An electronic device, characterized in that, Includes: A memory for storing a computer program; A processor for implementing the steps of the model loading method according to any one of claims 1 to 16 when executing the computer program.

19. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program, when executed by a processor, implements the steps of the model loading method according to any one of claims 1 to 16.

20. A computer program product, comprising a computer program, characterized in that, The computer program, when executed by a processor, implements the steps of the model loading method according to any one of claims 1 to 16.

Citation Information

Patent Citations

  • Object storage system and method used for cloud rendering

    CN108073350A

  • Model deployment method and device

    CN113407192A

  • Snapshot-based depth model rapid loading method and apparatus

    CN117270988A

  • Database management method, computer equipment, storage medium and program product

    CN118245471A

  • Large model service starting method and device, equipment and medium

    CN119376971A

Cited By

  • Task scheduling method, task scheduling device, electronic equipment and readable storage medium

    CN121364936A

  • A neural network model loading implementation method and device and medium

    CN122547416A

  • A data processing method and apparatus

    CN122572695A