Preheating method, device, equipment and system for model file

By controlling the nodes to obtain resource reporting information, automatically allocating target nodes and preheating model files, the problem of slow loading of large machine learning models is solved, and the response speed of model services and user experience are improved.

CN120596448APending Publication Date: 2025-09-05JINAN INSPUR DATA TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510740169.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

The loading and startup speed of large machine learning models in existing technologies is slow. Traditional loading methods are time-consuming and occupy system resources. In addition, there is a lack of automated model file preheating solutions, resulting in frequent preheating failures.

Method used

The preheating task scheduling management component is used to obtain resource reporting information through the control node, determine the target node based on the disk space information, and generate preheating task instructions. The file synchronization component is used to preheat the model file from the shared storage to the target node. It supports multiple storage types and performs retries and downgrades when the task fails.

Benefits of technology

It realizes the automatic preheating of model files, improves the response speed of model services and user experience, reduces human errors, and improves the efficiency and stability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120596448A_ABST
    Figure CN120596448A_ABST
Patent Text Reader

Abstract

The invention discloses a model file preheating method, device, equipment and system, and relates to the technical field of cloud computing, the method comprises the following steps: a control node uses a preheating task scheduling management component to obtain resource reporting information of each allocable node; determining a target node corresponding to the obtained preheating task request from the distributable nodes according to the resource reporting information, and generating a preheating task instruction corresponding to the target node; wherein the preheating task request comprises a model file storage path; sending the preheating task instruction to a target node; according to the method, the target node corresponding to the obtained preheating task request is determined from the distributable nodes according to the resource report information, so that the preheating task can be automatically distributed, the preheating task instruction is sent to the target node, the target node is controlled to cache the corresponding model file from the shared storage, and the preheating task can be automatically distributed. Automatic model file preheating is achieved, and the response speed of model service and the user experience are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of cloud computing technology, and in particular to a model file preheating method, device, equipment and system. Background Art

[0002] With the rapid development of artificial intelligence and machine learning technologies, the application of large machine learning models is becoming increasingly widespread. However, the loading and startup speed of large model files has become a major bottleneck. When deploying large models on cloud platforms, traditional loading methods typically require reading all model files from remote storage servers, which is not only time-consuming but also consumes a large amount of system resources. Currently, there are some optimization technologies for model file loading, such as distributed caching systems. However, in a distributed system, model files need to be transferred and loaded between multiple nodes, which not only increases loading time but also increases system complexity, making efficient file pre-warming difficult.

[0003] In related technologies, model file preheating relies primarily on manual configuration and scripting, lacking automated solutions. This approach is not only time-consuming but also prone to human error, leading to preheating failures. Furthermore, as model files continue to grow in size, these solutions are no longer able to meet practical needs. Therefore, how to automate model file preheating to improve the responsiveness and user experience of model services is an urgent issue. Summary of the Invention

[0004] The purpose of the present invention is to provide a model file preheating method, device, equipment and system to realize automatic model file preheating and improve the response speed of model service and user experience.

[0005] In order to solve the above technical problems, the present invention provides a model file preheating method, comprising:

[0006] The control node uses the preheating task scheduling management component to obtain resource reporting information of each assignable node; wherein the assignable node is a host node running the node resource management component, and the resource reporting information includes disk space information;

[0007] Determine, from the assignable nodes, a target node corresponding to the acquired warm-up task request according to the resource reporting information, and generate a warm-up task instruction corresponding to the target node; wherein the warm-up task request includes a model file storage path;

[0008] The preheating task instruction is sent to the target node, so that the target node uses the file synchronization component to preheat the model file of the model file storage path in the shared storage to the target node according to the preheating task instruction.

[0009] On the other hand, the preheating task request further includes a task priority, and determining a target node corresponding to the obtained preheating task request from the allocatable nodes according to the resource reporting information includes:

[0010] The target node corresponding to each of the preheating task requests is determined according to the available disk space in the disk space information, the task priority of each of the preheating task requests, and the pre-occupied space of each of the allocatable nodes.

[0011] On the other hand, after determining the target node corresponding to the obtained warm-up task request from the allocatable nodes according to the resource reporting information, the method further includes:

[0012] Controlling the target node to start the file synchronization component corresponding to the warm-up task request;

[0013] Correspondingly, sending the preheating task instruction to the target node includes:

[0014] The warm-up task instruction is sent to the file synchronization component of the target node.

[0015] In another aspect, the method further comprises:

[0016] The current target node uses the preheating task monitoring component to obtain task status information of the preheating task in the node; wherein the current target node is any of the target nodes, the preheating task in the node includes the preheating task executed by the file synchronization component in the current target node, and the task status information includes the current status;

[0017] The task status information is sent to the preheating task scheduling management component.

[0018] On the other hand, after sending the warm-up task instruction to the target node, the method further includes:

[0019] The control node uses the preheating task scheduling management component to determine the target preheating task instruction of the failed task according to the received task status information of each target node;

[0020] If the number of task failures of the current task instruction does not reach the number threshold, then control the target node corresponding to the current task instruction to retry the task; wherein the current task instruction is any of the target preheating task instructions;

[0021] If the number of task failures of the current task instruction does not reach the threshold number, the task priority of the current task instruction is lowered;

[0022] Correspondingly, sending the preheating task instruction to the target node includes:

[0023] Each of the preheating task instructions is sent to its corresponding target node according to its task priority.

[0024] In another aspect, the method further comprises:

[0025] The currently allocatable node utilizes the running node resource management component to predict the disk space usage of the currently allocatable node within a future preset time period based on the acquired historical disk space information; wherein the currently allocatable node is any of the aforementioned allocatable nodes;

[0026] If the disk space usage meets the prediction alarm condition, prediction alarm information is generated and sent to the warm-up task scheduling management component.

[0027] On the other hand, the preheating task instruction includes the model file storage path and the model file storage type; the method further includes:

[0028] The current target node utilizes the running file synchronization component to load the target storage plug-in according to the model file storage type in the received warm-up task instruction; wherein the target storage plug-in is any preset storage plug-in;

[0029] The target storage plug-in is used to preheat the target model file in the corresponding shared storage to the cache directory of the current target node.

[0030] The present invention also provides a model file preheating device, which is applied to a control node and includes:

[0031] A resource receiving module, configured to obtain resource reporting information of each allocatable node using the preheating task scheduling management component; wherein the allocatable node is a host node running the node resource management component, and the resource reporting information includes disk space information;

[0032] A task scheduling module, configured to determine, from the allocatable nodes, a target node corresponding to the acquired warm-up task request based on the resource reporting information, and generate a warm-up task instruction corresponding to the target node; wherein the warm-up task request includes a model file storage path;

[0033] The task issuing module is used to send the preheating task instruction to the target node, so that the target node uses the file synchronization component to preheat the model file of the model file storage path in the shared storage to the target node according to the preheating task instruction.

[0034] The present invention also provides a model file preheating device, comprising:

[0035] memory for storing computer programs;

[0036] A processor is used to implement the steps of the above-mentioned model file preheating method when executing the computer program.

[0037] In addition, the present invention also provides a model file preheating system, comprising: a first host device and a second host device communicatively connected to the first host device;

[0038] The first host device is a preheating device for the model file as described above; and the second host device runs a node resource management component.

[0039] The present invention provides a preheating method for a model file, comprising: a control node using a preheating task scheduling management component to obtain resource reporting information of each assignable node; wherein the assignable node is a host node running a node resource management component, and the resource reporting information includes disk space information; based on the resource reporting information, a target node corresponding to the obtained preheating task request is determined from the assignable node, and a preheating task instruction corresponding to the target node is generated; wherein the preheating task request includes a model file storage path; the preheating task instruction is sent to the target node, so that the target node uses a file synchronization component to preheat the model file in the model file storage path in the shared storage to the target node according to the preheating task instruction.

[0040] As can be seen, the present invention can automatically allocate preheating tasks by determining the target node corresponding to the preheating task request from the allocable nodes based on resource reporting information. This allows for automated model file preheating by sending preheating task instructions to the target node and controlling the target node to cache the corresponding model file from shared storage, thereby improving the response speed of the model service and the user experience. Furthermore, the present invention also provides a model file preheating device, equipment, and system, which also have the aforementioned beneficial effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0042] Figure 1 The present invention provides a flowchart of a method for preheating a model file.

[0043] Figure 2 A schematic diagram of the system architecture of another model file preheating method provided by an embodiment of the present invention;

[0044] Figure 3A structural block diagram of a model file preheating device provided by an embodiment of the present invention;

[0045] Figure 4 A schematic structural diagram of a model file preheating device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0046] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0047] Please refer to Figure 1 , Figure 1 This is a flow chart of a method for preheating a model file provided by an embodiment of the present invention. The method may include:

[0048] Step 101: The control node uses the preheating task scheduling management component to obtain resource reporting information of each allocatable node; wherein the allocatable node is a host node running the node resource management component, and the resource reporting information includes disk space information.

[0049] It can be understood that the control node in this embodiment is a host node running a preheating task scheduling management component, that is, a host node in a cluster (such as a distributed system) used to schedule and manage preheating tasks (that is, model file synchronization tasks); the assignable node in this embodiment can be a host node running a node resource management component, that is, a host node in a cluster used for scheduling and controlling controlled nodes, which can be assigned to preheating tasks and perform processing.

[0050] Accordingly, the specific selection of control nodes and allocable nodes in the cluster in this embodiment can be set by the designer based on practical scenarios and user needs. For example, allocable nodes can include all host nodes in the cluster, that is, the control node can also serve as an allocable node; allocable nodes can also include host nodes other than the control node in the cluster. This embodiment does not impose any restrictions on this.

[0051] Accordingly, the resource reporting information in this embodiment can be information collected and reported by each allocable node using its own running node resource management component; that is, the node resource management component can have resource monitoring (such as Figure 2 The disk monitoring function in the control node) and resource reporting function are used to report the monitored resource information (such as disk space information) to the warm-up task scheduling management component of the control node.

[0052] It should be noted that the specific content of the resource reporting information for each allocable node in this embodiment can be customized by the designer based on practical scenarios and user needs. For example, the resource reporting information may include disk space information, such as any one or more of the total space, used space, free space (i.e., available disk space), and utilization rate of each monitored disk partition within the allocation node. Resource reporting information may also include other information such as node identification and / or disk IO (input / output) metrics, which are not limited in this embodiment.

[0053] Correspondingly, the method provided in this embodiment may also include the current allocatable node utilizing the running node resource management component to obtain and send resource reporting information to the preheating task scheduling management component; wherein, the current allocatable node is any allocatable node, such as a control node serving as an allocatable node, or an allocatable node other than the control node.

[0054] Step 102: According to the resource reporting information, determine the target node corresponding to the obtained warm-up task request from the allocatable nodes, and generate a warm-up task instruction corresponding to the target node; wherein the warm-up task request includes the model file storage path.

[0055] It is understandable that in this step, the control node can use the preheating task scheduling management component to determine the target node corresponding to each preheating task request from all the assignable nodes based on the resource reporting information (such as disk space information) of each assignable node, so as to realize the scheduling function of the preheating task, such as Figure 2 The preheating task scheduling management component has a task scheduling function, and generates preheating task instructions corresponding to the target node to issue preheating tasks.

[0056] The preheating task request in this embodiment may be a request to preheat the model file (i.e., cache synchronization), such as a preheating task request for a large model file issued by a user configuration. The preheating task instruction in this embodiment may be an instruction issued to the target node for preheating the model file.

[0057] Correspondingly, the specific method of determining the target node corresponding to the obtained preheating task request from the allocatable nodes based on the resource reporting information in this step, that is, the specific task scheduling method of the preheating task scheduling management component, can be set by the designer according to the practical scenario and user needs. For example, the allocatable node with the largest available disk space can be directly determined as the target node corresponding to the current preheating task request; wherein, the current preheating task request is any preheating task request. When the preheating task request also includes a task priority, the target node corresponding to each preheating task request can be determined based on the available disk space in the disk space information, the task priority of each preheating task request, and the respective pre-occupied space (that is, the reserved disk space) of each allocatable node. This embodiment does not impose any restrictions on this.

[0058] Step 103: Send the preheating task instruction to the target node, so that the target node uses the file synchronization component to preheat the model file in the model file storage path in the shared storage to the target node according to the preheating task instruction.

[0059] In this step, the control node can utilize the preheating task scheduling management component to send each preheating task instruction to its corresponding target node, completing the dispatch of the preheating task. This enables each target node to utilize its own file synchronization component to preheat the model file located in the model file storage path in the shared storage to its own cache directory based on the received preheating task instruction. This embodiment does not limit the specific file content and type of the model file located in the model file storage path in the shared storage. For example, the model file can be a large model file, such as a large model weight file.

[0060] Correspondingly, the file synchronization component in this embodiment can be a component for caching model files from remote storage to a local cache directory on the host. For example, the file synchronization component can be started and run on all allocatable nodes, or it can be started and run only on allocatable nodes (i.e., target nodes) designated by the warm-up task scheduling management component. This embodiment does not impose any restrictions on this.

[0061] Furthermore, in this embodiment, a preheating task monitoring component can be started on the assignable node or the target node to monitor the preheating task in the node (i.e., the preheating task executed by the file synchronization component in the node).

[0062] The task status information is reported to the preheating task scheduling management component, so that the control node can understand the current status of each preheating task (such as running, completed, and failed).

[0063] It should be noted that the specific function settings and startup operation process of the above-mentioned file synchronization component, preheating task listening component, node resource management component and preheating task scheduling management component can be set by the designer according to the practical scenario and user needs. As long as the model file can be distributed and preheated from a shared storage, such as object storage or NFS (Network File System) storage, to the cache directory of a specific host node (i.e., the target node), this embodiment does not impose any restrictions on this.

[0064] Furthermore, since the preheating process of the model file involves the support of multiple storage types, and the related technologies often lack support for multiple storage types, resulting in limitations in the preheating process, the file synchronization component in this embodiment can support plug-in docking with different types of storage. For example, the current target node can use the running file synchronization component to load the target storage plug-in according to the model file storage type in the received preheating task instruction; use the target storage plug-in to preheat the target model file in the corresponding shared storage to the cache directory of the current target node; wherein the target storage plug-in is any preset storage plug-in, and the preset storage plug-in can include an NFS storage plug-in (such as Figure 2 NFS plugin in ) and object storage plugins, such as Figure 2 The S3 (Simple Storage Service, an object storage) plugin in .

[0065] For example, if Figure 2 As shown, the file synchronization component may include the following functions:

[0066] ①. Storage plug-in lifecycle management function: supports loading and unloading mechanisms for plug-ins of multiple storage types (i.e., preset storage plug-ins), and supports dynamic loading of storage plug-ins (i.e., target storage plug-ins) at runtime. Each storage plug-in can implement the following interfaces: 1) File list query interface, used to obtain file information that needs to be synchronized; 2) File download interface, used to download source files from shared storage to the cache directory of the host node; 3) File integrity verification interface, used to ensure that the downloaded file is consistent with the source file.

[0067] For example, during the startup phase, the file synchronization component first performs system initialization operations, including reading the configuration file, setting the log system and startup-related parameters. The configuration file defines various storage plug-in paths and basic download parameters (such as the default number of threads and default shard size), which will be used in the subsequent file synchronization process.

[0068] The file synchronization component's plugin loading mechanism scans the specified plugin directory based on the configuration file to identify available storage type plugins. Plugins are loaded into memory using a dynamic link library (DLL) and their initialization functions are called. During initialization, each plugin registers its supported storage types and implemented interface functions with the component, including file list query, file download, and file integrity verification interfaces.

[0069] The process of the file synchronization component receiving and analyzing the warm-up task instructions can receive the warm-up task instructions from the warm-up task scheduling management component. The instructions may include information such as the source storage type (i.e., the model file storage type), the source storage path (i.e., the model file storage path), the target local cache directory, the file list, and the task priority; parse the warm-up task instructions, extract various parameters, and verify whether the source storage type has corresponding plug-in support based on the loaded plug-in list; if there is no supporting plug-in, feedback error information to the scheduling management component and terminate the corresponding warm-up task.

[0070] ② Pre-download Verification: Before downloading, the resource reporting component checks the target node's disk space information. For example, the file synchronization component can call the disk space query interface provided by the node resource management component to obtain the current available disk space on the node (i.e., the target node). It can also query the source storage plug-in (i.e., the target storage plug-in) to obtain file size information and calculate the required space for the pre-loaded file. If the available disk space is less than the required space, the scheduling management component reports the insufficient space and terminates the task. If sufficient space is available, the subsequent operation continues.

[0071] ③. Smart Downloader Functions: 1) Supports adaptive segmented downloading (automatically enabled if the file size is >1GB); 2) Supports dynamic bandwidth adjustment (e.g. based on network quality detection); 3) Supports cross-node P2P (Peer-to-Peer, node-to-node) transmission; 4) Supports breakpoint resumption and multi-threaded downloading to improve the efficiency and stability of file synchronization.

[0072] For example, when the file synchronization component obtains a file list, it can use the source storage plug-in's file list query interface to obtain the list of files to be synchronized according to the source storage path specified by the task. For situations involving large numbers of files or complex directory structures, a recursive query method is used to ensure that complete file information is obtained.

[0073] The process of developing an intelligent download strategy in the file synchronization component may include: determining the size of a single file and initiating multi-segment download mode when the file size exceeds a segmentation threshold (e.g., 1GB); dynamically determining the number of segments based on factors such as network bandwidth, disk I / O performance, and file size, and appropriately allocating download threads to each segment to achieve adaptive segmented downloading. Segment sizes range from 64MB to 256MB. This may also include: initiating a network quality detection function to initially assess the current network bandwidth and determine an initial download speed limit; establishing a network status monitoring timer to periodically reassess the network bandwidth (e.g., every 30 seconds) during the download process, providing a basis for subsequent dynamic bandwidth adjustment. This may also include querying the local node's P2P transmission configuration and known information about other nodes participating in the warm-up task; if a node meets the requirements (e.g., located on the same local area network and with available bandwidth), establishing a P2P connection with it, negotiating the proportion of available P2P transmission bandwidth, and allocating the corresponding file segment download tasks.

[0074] The multi-threaded startup process of the file synchronization component includes: creating multiple download threads based on the specified download strategy; each thread is responsible for downloading a file segment or portion of the data. When the thread is started, the thread's download start and end locations, as well as related download parameters such as timeout and retry count, are set.

[0075] The breakpoint-resume download process of the file synchronization component may include: for files that have been partially downloaded before, checking whether the file exists in the local cache directory and its download progress; reading the existing download record file, obtaining the completed segment information or the download byte range, so that each thread can continue downloading from the last interrupted position.

[0076] The data download process of the file synchronization component can include: each download thread, based on its assigned task, reads data from remote storage through the source storage plug-in's file download interface and writes it to the corresponding file location in the local cache directory. During the download process, the thread monitors download speed, error count, and other information in real time.

[0077] The file synchronization component's dynamic bandwidth adjustment process can include: using network quality monitoring to periodically assess network bandwidth and compare it with the initial assessment results. If the current bandwidth drops below a threshold (e.g., 30%), the download speed limit for each thread is reduced accordingly; if the bandwidth increases, the speed limit can be increased accordingly to fully utilize network resources. Simultaneously, network bandwidth changes are fed back to the scheduling management component so it can adjust task scheduling.

[0078] The cross-node P2P transmission execution and monitoring process of the file synchronization component may include: if P2P transmission is enabled, real-time monitoring of the stability, transmission speed and data integrity of the P2P connection during the download process; when a P2P transmission node is found to be faulty or the transmission speed is too slow, adjusting the shard download task allocation to reduce dependence on the problem node.

[0079] The file integrity verification process of the file synchronization component may include: after each shard is downloaded, the shard data is verified using the file integrity verification interface provided by the source storage plug-in. For example, a verification method based on hash values ​​(such as the encryption algorithm MD5 or the secure hash algorithm SHA-1) is used to compare and verify the hash value generated by the downloaded shard data with the corresponding shard hash value provided by the source storage; if the shard verification fails, an error log is recorded and the shard is downloaded again. After all shards are downloaded and verified successfully, the shards are merged into a complete file, and the integrity of the merged file is again verified to ensure the correctness of the entire file; if the verification fails after merging, the cause of the failure is analyzed, such as incorrect shard order or damage to individual shard data, and appropriate repair measures are taken, such as re-downloading some shards or re-merging the file.

[0080] The file synchronization component may also include a task cleanup process, such as cleaning up related temporary files, download record files, and no longer needed shard files after the file synchronization is successfully completed to release occupied system resources, such as memory buffers and file handles.

[0081] The file synchronization component can also include a result feedback process, such as sending a task completion notification to the preheat task scheduling management component, which contains information such as the task ID (code), synchronization result (success or failure), and file storage path. It also records detailed execution logs for this task, including start time, end time, download speed, and any errors encountered and how they were handled, to facilitate subsequent task analysis and optimization.

[0082] Correspondingly, such as Figure 2 As shown, the preheating task monitoring component can include the following functions:

[0083] ①. Task status management function, which is used to maintain the task status information of the warm-up task, such as the task ID, current status and progress information. For example, when the warm-up task monitoring component is started, it can load the corresponding configuration file to initialize the log system, communication interface and internal data structure. The configuration file can include component operation parameters, such as the task scanning cycle, communication port and the address of the warm-up task scheduling management component. Create a task registry to store and manage relevant information of the file synchronization component running on the node, including key data such as task ID, start time, current status and progress information; the task registry can adopt a hash table structure with the task ID as the key to facilitate quick query and update of task status information.

[0084] The warmup task listener component periodically scans the process list on the host node and identifies running file synchronization components by their process name, port, or identifier. It also examines the startup parameters of these components and extracts key information, such as the task ID, to associate them with records in the task registry. For each detected file synchronization component, a dedicated listener (such as a thread or process) is established. This listener establishes a communication connection with the file synchronization component through inter-process communication mechanisms (such as signals, named pipes, and sockets) to obtain real-time execution status information. The listener periodically sends status query requests to the file synchronization component and waits for a response. Task status information obtained from the file synchronization component includes progress information (such as the number of completed files and the amount of data downloaded), current status (such as running, completed, or failed), error information (if any), and resource usage (such as CPU usage and memory usage).

[0085] The warm-up task monitoring component can also include a registry update process, such as updating the collected task status information to the corresponding task record in the task registry; if the current status of the warm-up task changes (such as from running to completed), the update time of the task record is marked and the status reporting process is triggered.

[0086] ②. Task reporting function, which is used to report task status information to the preheating task scheduling management component regularly or when the current status of the preheating task changes. That is, the method provided by this embodiment may also include: the current target node uses the preheating task monitoring component to obtain the task status information of the preheating task in the node; and sends the task status information to the preheating task scheduling management component. Among them, the preheating task in the node includes the preheating task executed by the file synchronization component in the current target node, and the task status information includes the current status; the current target node is any target node, such as a control node as a target node, or a target node other than a control node.

[0087] For example, the warm-up task monitoring component can send task status information to the warm-up task scheduling management component according to the reporting period in its configuration file (such as every 30 seconds) or the task status change; for warm-up tasks with high task priority and / or major changes in task status (such as failure or completion), the task status information is reported immediately; for regular status updates, the task status information is reported according to the reporting period.

[0088] The specific content of the task status information reported by the preheating task monitoring component can be set by the designer, such as extracting the task status information to be reported from the task registration table and organizing it into a prescribed reporting data format (i.e., a preset reporting data format), such as including fields such as task ID, current status, progress, node information (such as node identification), etc.; the task status information data is sent to the preheating task scheduling management component through a communication channel pre-established with the preheating task scheduling management component.

[0089] ③. Exception handling: This function is used to handle exceptions during the warm-up task execution. For example, the warm-up task monitoring component can monitor the operation of the file synchronization component. When it detects an abnormal situation such as the file synchronization component's abnormal exit, communication interruption, or task status information update timeout, it determines the type of exception. Examples of exception types include component crash, network failure resulting in inability to obtain status, and storage resource exhaustion.

[0090] The warmup task monitoring component can also take appropriate action based on the identified exception type. For example, if a component crashes, the file synchronization component can be restarted and the relevant error information can be logged. If a network failure occurs, a reconnection mechanism can be enabled to reestablish communication after the network is restored. If a storage resource issue occurs, the node resource management component can be notified for processing.

[0091] Correspondingly, such as Figure 2 As shown, the node resource management component can include the following functions:

[0092] ① Disk space monitoring function, used to regularly monitor a node's disk space usage (i.e., disk space information). For example, at startup, the node resource management component can set a disk monitoring period (e.g., every minute) based on its configuration file and create a monitoring timer. This configuration file can also include the disk partitions or mount points to be monitored, as well as warning thresholds (e.g., triggering a warning when free space falls below 20%). When the monitoring timer is triggered, the node resource management component can call the disk space query interface provided by the operating system. For example, in Linux (an operating system), by executing the "df -h" command or reading the " / proc / diskstats" file, disk space information such as the total space, used space, free space, and usage rate of each monitored disk partition can be obtained. The collected disk space information is stored in a local monitoring data storage structure; this structure can be a database, an in-memory hash table, or a dedicated monitoring data file. After each disk space information collection, the latest space information record for the corresponding disk partition is updated, and the data collection timestamp is recorded.

[0093] ②. Resource reporting function, which is used to periodically send disk space information to the task scheduling management component; and when there is a significant change in disk space information, it is reported immediately to inform the task scheduling management component of the latest disk space status in a timely manner.

[0094] For example, the node resource management component can extract the disk space information that needs to be reported from the local monitoring data storage structure, including the total space, available space, usage rate, and node identification of each monitoring partition. The disk space information that needs to be reported is organized according to a pre-set reporting data format, such as JSON (an open standard file format and data exchange format), to ensure the integrity and accuracy of the data. The node resource management component can pre-establish a communication connection with the "Warm-up Task Scheduling Management Component"; during the reporting process, the status of the communication connection is continuously monitored; if the connection is interrupted, an attempt is made to re-establish the connection, and the unsuccessful reported data is cached, waiting for a successful reconnection before re-reporting.

[0095] The node resource management component can send disk space information to the warmup task scheduling component at a preset resource reporting interval (e.g., every 5 minutes). Furthermore, if there are significant changes in disk space information (e.g., a decrease in available space exceeding an alarm threshold, such as 5GB), the reporting process is immediately triggered, promptly notifying the warmup task scheduling component of the latest disk space status.

[0096] ③. The disk space warning function is used to predict disk space and trigger the warning mechanism when the warning conditions are met according to the preset warning rules. For example, the node resource management component can compare the current available space with the preset warning threshold after each update of the disk space information to determine whether the warning conditions are met. The warning threshold can be a fixed value (such as 100GB), for example, it is determined that a certain warning condition is met when the available space is less than 100GB; the warning threshold can also be a relative value (such as 80%), for example, it is determined that a certain warning condition is met when the usage rate exceeds 80%. The node resource management component can generate corresponding warning information when a certain warning condition is met, such as node information, disk partition information, current available space and warning level; at the same time, it can record the warning event locally, including detailed information such as warning time and cause, for subsequent analysis and auditing.

[0097] The node resource management component can also predict disk space usage within a preset time period in the future based on the historical disk space information obtained. For example, the node resource management component can collect disk space information (i.e., historical disk space information) within a preset collection time period (e.g., 1 hour), such as total space, used space, and available space. It can perform data preprocessing on the historical disk space information, such as removing outliers, interpolating missing data, and normalizing it to the range of [0,1]. It can also construct an LSTM (Long Short-Term Memory) model, divide the training set and test set into chronological order, select the mean square error (MSE) as the loss function, use the Adam (Adaptive Moment Estimation, an optimization algorithm) optimizer for gradient descent, and use the trained LSTM model to predict disk space usage within a preset time period in the future (e.g., 2 hours). That is to say, the method provided in this embodiment may also include: the current allocatable node uses the running node resource management component to predict the disk space usage of the current allocatable node within a future preset time period based on the acquired historical disk space information; if the disk space usage meets the prediction alarm conditions, a prediction alarm information is generated and sent to the preheating task scheduling management component; wherein the current allocatable node is any allocatable node.

[0098] Correspondingly, the currently allocatable node can also utilize the running node resource management component to generate and send prediction alarm information to the preheating task scheduling management component when the disk space usage meets the prediction alarm conditions (such as when the available space in the disk space usage is less than the prediction alarm threshold).

[0099] ④. Dynamic disk space allocation function, used to dynamically adjust the disk space of the cache directory based on task requirements. For example, upon receiving an expansion instruction from the preheating task scheduling management component or determining that expansion is needed based on forecasted alarm information and / or early warning information, the node resource management component can determine the required expansion space based on historical disk space usage trends based on historical disk space information (or predicted disk space usage within a preset future time period) and the storage requirements of the preheating task. The determination of expansion space can also take into account reserving a certain amount of buffer space (such as 1.2 times the estimated demand).

[0100] Accordingly, the node resource management component can expand capacity based on the node's storage architecture and available capacity, adopting a corresponding expansion solution based on the node's storage architecture. For example, if the node's storage architecture uses LVM (Logical Volume Manager), you can increase the size of the logical volume by calling the LVM command-line tool.

[0101] Correspondingly, the node resource management component can also provide feedback on expansion results (such as expansion success or failure) to the warmup task scheduling management component so that it can adjust its task scheduling strategy. It also locally records detailed information about the expansion operation, including expansion time, expansion size, and error messages during the operation, to facilitate subsequent space management analysis and troubleshooting.

[0102] ⑤. Disk I / O performance monitoring: This function monitors disk I / O performance to prevent task delays caused by I / O bottlenecks. For example, the node resource management component can utilize the disk I / O performance monitoring interface provided by the operating system, such as the " / proc / diskstats" file or iostat (a tool used to monitor system input / output device load) in Linux, to regularly collect disk I / O performance metrics such as read / write speed, read / write latency, and queue length. The disk monitoring period can be short (e.g., every 10 seconds) to promptly capture fluctuations in disk I / O performance.

[0103] Correspondingly, such as Figure 2 As shown, the preheating task scheduling management component may include the following functions:

[0104] ①. Task scheduling function, used to reasonably schedule warm-up tasks based on the resource reporting information of each allocatable node. For example, the warm-up task scheduling management component can receive warm-up task requests for model files issued by users through an interface that interacts with users (such as the application programming interface RestAPI); the warm-up task request can include parameter information such as the model file storage path, target node range, and task priority. Parse the warm-up task request and extract the parameter information in the request; based on the parameter information, verify the legitimacy of the warm-up task request, such as checking whether the model file storage path is valid, whether the target node range complies with the cluster configuration, and whether the task priority is within the allowed range; for illegal warm-up task requests, return an error message to the user and reject the task submission.

[0105] The warm-up task scheduling management component obtains disk space information for allocatable nodes from the node resource management component. This information is reported periodically by the node resource management component and retrieved on-demand during task scheduling. Based on the requirements of legitimate warm-up task requests (required disk space) and the disk space information of each node, available nodes that meet the basic running conditions of the task are selected. Unavailable nodes with limited resources are filtered out based on preset resource usage thresholds (e.g., disk space usage below 30%).

[0106] The warmup task scheduling management component can adopt a hybrid scheduling strategy, taking into account factors such as available disk space (prioritizing allocation to nodes with sufficient space), task priority (high-priority tasks receive priority resource allocation), and pre-occupied space (reserving node resources in advance for potential subsequent large tasks). This strategy determines the allocation ratio and launch order of each task on each node, resulting in a task allocation plan. According to the task allocation plan, a task launch instruction is sent to the selected node (i.e., the target node). The instruction may include information such as the model file storage path, the target cache directory, and the task ID. Reliable communication protocols (such as remote procedure calls (RPCs)) are used to ensure that the task launch instruction is accurately transmitted to the file synchronization component on the node.

[0107] Correspondingly, in this embodiment, before sending the preheating task instruction to the target node, it is necessary to start the corresponding file synchronization component at the target node; as in step 102, after determining the target node corresponding to the obtained preheating task request from the allocable nodes based on the resource reporting information, the target node can be controlled to start the file synchronization component corresponding to the preheating task request, so that the preheating task instruction can be sent to the file synchronization component of the target node in step 103.

[0108] ②. Task fault tolerance and degradation function, used to support automatic retry and degradation after task failure. For example, after the preheating task instruction is sent to the target node in step 103, the control node can use the preheating task scheduling management component to determine the target preheating task instruction for the task failure based on the task status information of each target node received; if the number of task failures of the current task instruction does not reach the number threshold, the target node corresponding to the current task instruction is controlled to retry the task; wherein, the current task instruction is any target preheating task instruction; if the number of task failures of the current task instruction does not reach the number threshold, the task priority of the current task instruction is reduced. Accordingly, the process of sending the preheating task instruction to the target node can send each preheating task instruction to its corresponding target node according to the task priority of each preheating task instruction.

[0109] For example, the preheating task scheduling management component can receive task status information reported by the preheating task listening component, including task ID, progress information, current status, and encountered errors. This task status information is categorized and organized, and stored in a task status database or in-memory status cache structure according to task ID. For automatically recoverable failed tasks (i.e., preheating tasks currently in a failed state), such as file synchronization interruptions caused by network fluctuations, the preheating task scheduling management component can automatically restart the failed portion of the task based on the task's retry strategy. During the retry process, the retry interval is gradually increased (e.g., with an exponential backoff strategy) to avoid excessive impact on system resources. Detailed information about each retry, including the retry time and failure reason, is recorded for subsequent analysis. When the number of task failures exceeds a threshold (e.g., 5) or it is determined that the preheating task cannot be successfully executed under current conditions, task downgrade is initiated, reducing the preheating task's performance. For example, based on the preheating task's importance and urgency, a decision can be made to downgrade the task to a suspended state or terminate it entirely.

[0110] In this embodiment, the embodiment of the present invention can realize automatic allocation of preheating tasks by determining the target node corresponding to the obtained preheating task request from the allocatable nodes based on resource reporting information, thereby sending the preheating task instruction to the target node and controlling the target node to cache the corresponding model file from the shared storage, thereby realizing automatic model file preheating and improving the response speed of the model service and user experience.

[0111] Corresponding to the above method embodiment, an embodiment of the present invention further provides a model file preheating device. The model file preheating device described below and the model file preheating method described above can refer to each other.

[0112] Please refer to Figure 3 , Figure 3This is a structural block diagram of a model file preheating device provided by an embodiment of the present invention. The device is applied to a control node and may include:

[0113] The resource receiving module 10 is used to obtain resource reporting information of each allocatable node using the preheating task scheduling management component; wherein the allocatable node is a host node running the node resource management component, and the resource reporting information includes disk space information;

[0114] The task scheduling module 20 is used to determine the target node corresponding to the obtained warm-up task request from the assignable nodes based on the resource reporting information, and generate a warm-up task instruction corresponding to the target node; wherein the warm-up task request includes the model file storage path;

[0115] The task issuing module 30 is used to send the preheating task instruction to the target node, so that the target node uses the file synchronization component to preheat the model file in the model file storage path in the shared storage to the target node according to the preheating task instruction.

[0116] On the other hand, the warm-up task request further includes a task priority, and the task scheduling module 20 may include:

[0117] The allocation submodule is used to determine the target node corresponding to each preheating task request according to the available disk space in the disk space information, the task priority of each preheating task request and the pre-occupied space of each allocatable node.

[0118] In another aspect, the task scheduling module 20 may further include:

[0119] The component control submodule is used to control the target node to start the file synchronization component corresponding to the preheating task request after determining the target node corresponding to the preheating task request from the allocable nodes according to the resource reporting information;

[0120] Correspondingly, the task issuing module 30 may be specifically configured to send the preheating task instruction to the file synchronization component of the target node.

[0121] In another aspect, the apparatus is applied to the current target node and may further include:

[0122] A status acquisition module is used to obtain task status information of the preheating task in the node using the preheating task monitoring component; wherein the current target node is any target node, the preheating task in the node includes the preheating task executed by the file synchronization component in the current target node, and the task status information includes the current status;

[0123] The status reporting module is used to send task status information to the preheating task scheduling management component.

[0124] In another aspect, the apparatus may further comprise:

[0125] A failure determination module is used to determine the target preheating task instruction that failed the task based on the received task status information of each target node using the preheating task scheduling management component;

[0126] A retry module is used to control the target node corresponding to the current task instruction to retry the task if the number of task failures of the current task instruction does not reach a threshold number; wherein the current task instruction is any target preheating task instruction;

[0127] A demotion module is used to reduce the task priority of the current task instruction if the number of task failures of the current task instruction does not reach a number threshold;

[0128] Correspondingly, the task issuing module 30 may be specifically configured to send each preheating task instruction to its corresponding target node according to the task priority of each preheating task instruction.

[0129] In another aspect, the apparatus is applied to a currently allocatable node and may further include:

[0130] A space prediction module is used to use the running node resource management component to predict the disk space usage of the current allocatable node within a preset time period in the future based on the acquired historical disk space information; wherein the current allocatable node is any allocatable node;

[0131] The prediction alarm module is used to generate and send prediction alarm information to the preheating task scheduling management component if the disk space usage meets the prediction alarm conditions.

[0132] On the other hand, the warm-up task instruction includes a model file storage path and a model file storage type; the device is applied to the currently allocable node and may further include:

[0133] A plug-in loading module is used to utilize the running file synchronization component to load a target storage plug-in according to the model file storage type in the received warm-up task instruction; wherein the target storage plug-in is any preset storage plug-in;

[0134] The file preheating module is used to use the target storage plug-in to preheat the target model file in the corresponding shared storage to the cache directory of the current target node.

[0135] In this embodiment, the embodiment of the present invention determines the target node corresponding to the obtained preheating task request from the allocatable nodes based on the resource reporting information through the task scheduling module 20, and can realize automatic allocation of the preheating task, thereby sending the preheating task instruction to the target node and controlling the target node to cache the corresponding model file from the shared storage, thereby realizing automatic model file preheating, and improving the response speed of the model service and user experience.

[0136] Corresponding to the above method embodiment, an embodiment of the present invention further provides a model file preheating device. The model file preheating device described below and the model file preheating method described above can refer to each other.

[0137] Please refer to Figure 4 , Figure 4 This is a schematic diagram of the structure of a model file preheating device provided by an embodiment of the present invention. The preheating device may include:

[0138] Memory D1, for storing computer programs;

[0139] The processor D2 is configured to implement the steps of the model file preheating method provided in the above method embodiment when executing a computer program.

[0140] The preheating device for the model file provided in this embodiment may specifically be a host device, such as the control node running the preheating task scheduling management component in the above embodiment.

[0141] Corresponding to the above device embodiment, an embodiment of the present invention further provides a model file preheating system. The model file preheating system described below and the model file preheating device described above can refer to each other.

[0142] A model file preheating system includes: a first host device and a second host device communicatively connected to the first host device;

[0143] Among them, the first host device is a preheating device of the model file provided in the above-mentioned device embodiment; the second host device runs a node resource management component; that is, the first host device in this embodiment can be a host device in the cluster that runs a preheating task scheduling management component, and the second host device in this embodiment can be a host device other than the first host device in the cluster that runs a node resource management component.

[0144] Corresponding to the above method embodiment, an embodiment of the present invention further provides a computer program product. The computer program product described below and the preheating method of a model file described above can refer to each other.

[0145] A computer program product includes a computer program / instruction, which, when executed by a processor, implements the steps of the model file preheating method provided in the above method embodiment.

[0146] Corresponding to the above method embodiment, an embodiment of the present invention further provides a computer-readable storage medium. The computer-readable storage medium described below and the preheating method of a model file described above can refer to each other.

[0147] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the model file preheating method of the above method embodiment.

[0148] The computer-readable storage medium may be a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, which may store program codes.

[0149] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. References to the common and similar parts between the various embodiments are sufficient. The devices, equipment, systems, computer program products, and computer-readable storage media disclosed in the embodiments are described briefly because they correspond to the methods disclosed in the embodiments. For relevant details, refer to the description of the methods.

[0150] The above is a detailed introduction to the preheating method, device, equipment and system of a model file provided by the present invention. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present invention, the present invention can also be improved and modified in several ways, and these improvements and modifications also fall within the scope of protection of the present invention.

Claims

1. A preheating method for a model file, characterized in that: include: The control node uses the preheating task scheduling management component to obtain resource reporting information of each assignable node; wherein the assignable node is a host node running the node resource management component, and the resource reporting information includes disk space information; Determine, from the assignable nodes, a target node corresponding to the acquired warm-up task request according to the resource reporting information, and generate a warm-up task instruction corresponding to the target node; wherein the warm-up task request includes a model file storage path; The preheating task instruction is sent to the target node, so that the target node uses the file synchronization component to preheat the model file of the model file storage path in the shared storage to the target node according to the preheating task instruction.

2. The preheating method of the model file according to claim 1, characterized in that: The preheating task request further includes a task priority. Determining a target node corresponding to the obtained preheating task request from the allocatable nodes based on the resource reporting information includes: The target node corresponding to each of the preheating task requests is determined according to the available disk space in the disk space information, the task priority of each of the preheating task requests, and the pre-occupied space of each of the allocatable nodes.

3. The preheating method of the model file according to claim 1, characterized in that: After determining the target node corresponding to the obtained warm-up task request from the allocatable nodes according to the resource reporting information, the method further includes: Controlling the target node to start the file synchronization component corresponding to the warm-up task request; Correspondingly, sending the preheating task instruction to the target node includes: The warm-up task instruction is sent to the file synchronization component of the target node.

4. The preheating method of the model file according to claim 1, characterized in that: Also includes: The current target node uses the preheating task monitoring component to obtain task status information of the preheating task in the node; wherein the current target node is any of the target nodes, the preheating task in the node includes the preheating task executed by the file synchronization component in the current target node, and the task status information includes the current status; The task status information is sent to the preheating task scheduling management component.

5. The preheating method of the model file according to claim 4, characterized in that: After sending the warm-up task instruction to the target node, the method further includes: The control node uses the preheating task scheduling management component to determine the target preheating task instruction of the failed task according to the received task status information of each target node; If the number of task failures of the current task instruction does not reach the number threshold, then control the target node corresponding to the current task instruction to retry the task; wherein the current task instruction is any of the target preheating task instructions; If the number of task failures of the current task instruction does not reach the threshold number, the task priority of the current task instruction is lowered; Correspondingly, sending the preheating task instruction to the target node includes: Each of the preheating task instructions is sent to its corresponding target node according to its task priority.

6. The preheating method of the model file according to claim 4, characterized in that: Also includes: The currently allocatable node utilizes the running node resource management component to predict the disk space usage of the currently allocatable node within a future preset time period based on the acquired historical disk space information; wherein the currently allocatable node is any of the aforementioned allocatable nodes; If the disk space usage meets the prediction alarm condition, prediction alarm information is generated and sent to the warm-up task scheduling management component.

7. The method for preheating a model file according to any one of claims 1 to 6, characterized in that: The preheating task instruction includes the model file storage path and the model file storage type; the method further includes: The current target node utilizes the running file synchronization component to load the target storage plug-in according to the model file storage type in the received warm-up task instruction; wherein the target storage plug-in is any preset storage plug-in; The target storage plug-in is used to preheat the target model file in the corresponding shared storage to the cache directory of the current target node.

8. A model file preheating device, characterized in that: Applicable to control nodes, including: A resource receiving module, configured to obtain resource reporting information of each allocatable node using the preheating task scheduling management component; wherein the allocatable node is a host node running the node resource management component, and the resource reporting information includes disk space information; A task scheduling module, configured to determine, from the allocatable nodes, a target node corresponding to the acquired warm-up task request based on the resource reporting information, and generate a warm-up task instruction corresponding to the target node; wherein the warm-up task request includes a model file storage path; The task issuing module is used to send the preheating task instruction to the target node, so that the target node uses the file synchronization component to preheat the model file of the model file storage path in the shared storage to the target node according to the preheating task instruction.

9. A model file preheating device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the model file preheating method according to any one of claims 1 to 7 when executing the computer program.

10. A model file preheating system, characterized in that: include: a first host device and a second host device communicatively connected to the first host device; Wherein, the first host device is a preheating device for the model file according to claim 9; and the second host device runs a node resource management component.

Citation Information

Patent Citations

  • Node scheduling method applied to distributed system and related device

    CN118152103A

  • KServe model reasoning acceleration method and device based on distributed cache and medium

    CN118153694A

  • FaaS application cold start acceleration method and device based on mirror image preheating

    CN118819669A

  • Server-free intelligent resource scheduling system based on large language model

    CN119883626A

  • Resource preheating method and related device

    CN120021235A