Resource preloading method and device for distributed computing, computing equipment and medium
By creating a cluster in distributed computing and using resource preloading and hash value verification methods, the problem of low resource preloading efficiency is solved, and fast task execution and efficient resource management are realized, which is suitable for distributed computing environments.
Patent Information
- Application Number
- CN202510421084.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-18
AI Technical Summary
How to efficiently and conveniently implement resource preloading of distributed computing to reduce the startup delay of task execution and improve computing efficiency.
Create a cluster containing head nodes and work nodes. Create processes in work nodes based on the process parameters of the calculation task through the head node, and use the resource download address to perform resource preloading, combine hash verification to ensure resource correctness, and dynamically manage processes when resource updates.
It significantly reduces the startup delay of task execution, improves the efficiency of distributed computing, ensures resource correctness, and supports resource updates and process management.
Smart Images

Figure CN120335887A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of Internet technologies, and more particularly, to a method, apparatus, computing device, and medium for preloading resources in distributed computing. Background Art
[0002] Distributed computing is a computing method in which each computing task is dispersed to multiple computing nodes for collaborative completion. Through distributed computing, the computing resources of multiple nodes can be effectively utilized to accelerate task processing and improve computing efficiency, scalability, and fault tolerance. In a distributed computing environment, each computing task often requires a large amount of resources such as data. For example, the resources may specifically include AI models, training data sets, pictures, etc. Preloading the resources required during the execution of a computing task can effectively improve the efficiency of distributed computing. Specifically, before the execution of a computing task, by preloading the resources required for the computing task into the computing nodes in advance, the startup latency of task execution can be significantly reduced. Therefore, how to efficiently and conveniently achieve resource preloading in distributed computing has become an urgent problem to be solved. Summary of the Invention
[0003] In view of the above problems, the present application proposes a method, apparatus, computing device, and medium for preloading resources in distributed computing to solve the following problem: how to efficiently and conveniently achieve resource preloading in distributed computing.
[0004] According to one aspect of the embodiments of the present application, a method for preloading resources in distributed computing is provided, including:
[0005] Create a cluster including a head node and worker nodes;
[0006] The head node creates a process corresponding to the computing task in the worker nodes according to the process parameters preset for the computing task;
[0007] Download the corresponding resources for the process according to the resource download address in the process parameters, and preload the resources;
[0008] Receive a computing task sent by a requester, and the head node submits the computing task to the process corresponding to the computing task for execution.
[0009] Optionally, creating a cluster including a head node and worker nodes further includes:
[0010] Create the head node and worker nodes through a cloud platform to form a cluster.
[0011] Optionally, downloading the corresponding resources for the process according to the resource download address in the process parameters, and preloading the resources further includes:
[0012] Download the resources required by the process during the execution of the computing task to the working node where the process is located according to the resource download address in the process parameters, and verify the resources;
[0013] If the verification passes, it is determined that the process is successfully created. Instantiate the singleton class corresponding to the resource according to the preloaded resource class path and resource class name in the process parameters, and preload the resources;
[0014] If the verification fails, return a notification that the process creation fails.
[0015] Optionally, further verifying the resources includes:
[0016] Use a preset hash algorithm to calculate the current hash value of the resource;
[0017] Compare whether the current hash value is the same as the target hash value of the resource in the process parameters;
[0018] If the current hash value is the same as the target hash value, it is determined that the verification passes;
[0019] If the current hash value is not the same as the target hash value, it is determined that the verification fails.
[0020] Optionally, the method further includes:
[0021] After the process is successfully created, add the process reference corresponding to the process to the process reference queue corresponding to the computing task, and manage the process reference queue through the head node.
[0022] Optionally, the head node submitting the computing task to the process corresponding to the computing task for execution further includes:
[0023] The head node selects a target process reference from the process reference queue corresponding to the computing task, and submits the computing task to the process corresponding to the target process reference in the working node for execution through the target process reference.
[0024] Optionally, the method further includes:
[0025] Manage the life cycle of the process corresponding to the computing task.
[0026] Optionally, the method further includes:
[0027] When the resource download address in the process parameters is updated, the head node creates a new process corresponding to the computing task in the working node according to the process parameters;
[0028] Download the corresponding updated resources for the new process according to the updated resource download address in the process parameters, and preload the updated resources;
[0029] The head node destroys the original process corresponding to the computing task.
[0030] Optionally, the method further includes:
[0031] Adjust the number of worker nodes in the cluster according to the usage of the cluster's hardware resources.
[0032] According to another aspect of the embodiments of the present application, a resource preloading device for distributed computing is provided, including:
[0033] A cluster creation module, adapted to create a cluster including a head node and worker nodes;
[0034] A process processing module, adapted to create a process corresponding to the computing task in the worker nodes by the head node according to the process parameters preset for the computing task;
[0035] A preloading module, adapted to download the corresponding resources for the process according to the resource download address in the process parameters and preload the resources;
[0036] A task execution module, adapted to receive the computing task sent by the requester, and the head node submits the computing task to the process corresponding to the computing task for execution.
[0037] Optionally, the cluster creation module is further adapted to:
[0038] Create a head node and worker nodes through a cloud platform to form a cluster.
[0039] Optionally, the preloading module is further adapted to:
[0040] Download the resources required by the process during the execution of the computing task to the worker node where the process is located according to the resource download address in the process parameters, and verify the resources;
[0041] If the verification passes, it is determined that the process creation is successful, and the singleton class corresponding to the resource is instantiated according to the preloaded resource class path and resource class name in the process parameters, and the resources are preloaded;
[0042] If the verification fails, a process creation failure notification is returned.
[0043] Optionally, the preloading module is further adapted to:
[0044] Use a preset hash algorithm to calculate the current hash value of the resource;
[0045] Compare whether the current hash value is consistent with the target hash value of the resource in the process parameters;
[0046] If the current hash value is consistent with the target hash value, it is determined that the verification passes;
[0047] If the current hash value is inconsistent with the target hash value, it is determined that the verification fails.
[0048] Optionally, the process processing module is further adapted to:
[0049] After the process is successfully created, add the process reference corresponding to the process to the process reference queue corresponding to the computing task, and manage the process reference queue through the head node.
[0050] Optionally, the task execution module is further adapted to:
[0051] The head node selects a target process reference from the process reference queue corresponding to the computing task, and submits the computing task to the process corresponding to the target process reference in the worker node for execution through the target process reference.
[0052] Optionally, the process processing module is further adapted to:
[0053] Manage the life cycle of the process corresponding to the computing task.
[0054] Optionally, the process processing module is further adapted to: when there is an update in the resource download address in the process parameters, the head node creates a new process corresponding to the computing task in the worker node according to the process parameters;
[0055] The preloading module is further adapted to: download the corresponding updated resources for the new process according to the updated resource download address in the process parameters, and preload the updated resources;
[0056] The process processing module is further adapted to: the head node destroys the original process corresponding to the computing task.
[0057] Optionally, the apparatus is further adapted to:
[0058] Adjust the number of worker nodes in the cluster according to the cluster hardware resource usage of the cluster.
[0059] According to another aspect of the embodiments of the present application, a computing device is provided, including: a processor, a memory, a communication interface, and a communication bus, and the processor, the memory, and the communication interface complete communication with each other through the communication bus;
[0060] The memory is used to store at least one executable instruction, and the executable instruction causes the processor to execute the operations corresponding to the above-mentioned resource preloading method for distributed computing.
[0061] According to still another aspect of the embodiments of the present application, a computer storage medium is provided, and at least one executable instruction is stored in the storage medium, and the executable instruction causes the processor to execute the operations corresponding to the above-mentioned resource preloading method for distributed computing.
[0062] In another aspect of the embodiments of the present application, there is provided a computer program product including at least one executable instruction, which causes a processor to perform operations corresponding to the above-mentioned resource preloading method for distributed computing.
[0063] According to the resource preloading method, device, computing device, medium and product for distributed computing provided by the embodiments of the present application, a cluster including a head node and worker nodes is created. The head node creates a process corresponding to the computing task in the worker nodes according to the process parameters corresponding to the computing task, downloads the corresponding resources for the process according to the resource download address in the process parameters, and preloads the resources, efficiently and conveniently realizing the resource preloading for distributed computing, enabling the process to quickly utilize the preloaded resources to execute the computing task, significantly reducing the startup latency of task execution, and effectively improving the efficiency of distributed computing; after the resource download is completed, the downloaded resources are verified in combination with the hash value of the resources, effectively ensuring that the downloaded resources are the resources required for executing the computing task, and helping to avoid the situation where the computing task execution goes wrong due to incorrect resource download; this solution not only realizes resource preloading, but also the resource class implements the singleton pattern, instantiating the singleton class corresponding to the resource when the process is created, ensuring that there is only one instance of the singleton class corresponding to the resource during the entire program running process, and providing a global access point to obtain this instance, conveniently keeping the resource resident; in addition, this solution can also conveniently handle the situation where the resource is updated. When the resource is updated, the head node creates a new process corresponding to the computing task and destroys the original process corresponding to the computing task, realizing the effective management of the process.
[0064] The above description is only an overview of the technical solutions of the embodiments of the present application. In order to be able to understand the technical means of the embodiments of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the embodiments of the present application more obvious and understandable, the following specifically describes the specific embodiments of the present application. Description of the Drawings
[0065] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the embodiments of the present application. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0066] Figure 1 A flowchart showing the resource preloading method for distributed computing according to an embodiment of the present application is shown;
[0067] Figure 2aShows a schematic flowchart of a resource preloading method for distributed computing according to another embodiment of the present application;
[0068] Figure 2b Shows a schematic architecture diagram of a resource preloading method for distributed computing;
[0069] Figure 3 Shows a block diagram of the structure of a resource preloading device for distributed computing according to an embodiment of the present application;
[0070] Figure 4 Shows a schematic diagram of the structure of a computing device according to an embodiment of the present application. Detailed implementation manners
[0071] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.
[0072] First, the noun terms involved in one or more embodiments of the present application are explained.
[0073] Distributed computing framework: Used to expand application programs, provides a computing layer for parallel processing, has efficient task parallelization and distributed computing capabilities, and is applicable to fields such as machine learning, data processing, and real-time systems.
[0074] Process: A concept in the distributed computing framework, is a stateful computing unit that allows parallel and distributed execution of computing tasks.
[0075] Cloud platform: Refers to a cloud computing platform based on Kubernetes (abbreviated as K8S), used for automated deployment, scaling, and management of containerized application programs, and is one of the key technical components for realizing agile development, rapid iteration, resource optimization, and flexible scaling.
[0076] Singleton: A software design pattern that ensures that there is only one instance of a certain class throughout the program's running process and provides a global access point.
[0077] Process reference: Another concept in the distributed computing framework, is a reference to a process that has been created in the distributed computing architecture. When a process is created, the distributed computing architecture can return a process reference, through which users can interact with the remote process, call its methods, and obtain results.
[0078] Lifecycle: It refers to the various stages experienced by an object, process, task, or system, etc., from creation, use to destruction.
[0079] Figure 1 The flowchart shows a resource preloading method for distributed computing according to an embodiment of the present application, as Figure 1 shown, the method includes the following steps:
[0080] Step S101, create a cluster including a head node and worker nodes.
[0081] Among them, a large amount of data and other resources are usually required during the execution of a computing task. The embodiment of the present application provides a solution for preloading the resources required for distributed computing in a distributed computing framework. Among them, the preloaded resources refer to the data resources required during the execution of the computing task, rather than the hardware resources such as CPU, GPU, and memory required for the execution of the computing task.
[0082] In a distributed computing framework, a process can be used to execute a computing task. In the embodiment of the present application, before the process executes the computing task, the resources required by the process during the execution of the computing task are preloaded, thereby significantly reducing the startup latency of task execution and effectively improving the efficiency of distributed computing.
[0083] In step S101, a cluster can be created through a cloud platform. Among them, the cluster includes a head node and worker nodes. Specifically, the number of head nodes can be one, and the number of worker nodes can be multiple; the head node serves as the control center of the cluster, responsible for scheduling tasks, managing, and maintaining the global state, etc., and the worker nodes are mainly responsible for executing actual computing tasks, receiving task instructions from the head node, executing the computing tasks, and reporting the execution results.
[0084] Step S102, the head node creates a process corresponding to the computing task in the worker nodes according to the process parameters corresponding to the computing task set in advance.
[0085] For the convenience of resource preloading and computing task execution, process parameters corresponding to the computing task are set in advance according to the task situation of the computing task, etc. The process parameters refer to the relevant parameters required for process creation and process execution. Among them. The process parameters may include parameters such as the number of replicas, hardware resource quotas, runtime environment variables, working directories, dependent data, resource download addresses, etc.
[0086] In the embodiment of the present application, the head node in the cluster is responsible for creating a process corresponding to the computing task. Specifically, the head node can create a process corresponding to the computing task in the worker nodes according to the number of replicas, hardware resource quotas, runtime environment variables, working directories, and dependent data in the process parameters.
[0087] Step S103: Download the corresponding resources for the process according to the resource download address in the process parameters, and preload the resources.
[0088] Among them, the resource download address refers to the address used to download the resources required during the execution of the computing task. Download the corresponding resources for the process according to the resource download address, and preload the resources, thereby realizing the resource preloading of distributed computing.
[0089] The resources may specifically include AI models, training data sets, pictures, etc. The resources required for different computing tasks are different. For example, if computing task 1 is to use a watermark removal intelligent model to remove the watermark in the to-be-processed picture, then the resources required for computing task 1 may include the watermark removal intelligent model. For the process created corresponding to computing task 1, the watermark removal intelligent model needs to be preloaded; another example is that if computing task 2 is to add an animation effect to the to-be-processed video, then the resources required for computing task 2 may include the animation effect-related data. For the process created corresponding to computing task 2, the animation effect-related data needs to be preloaded; yet another example is that if computing task 3 is to use a video copywriting generation model to generate video copywriting, then the resources required for computing task 3 may include the video copywriting generation model. For the process created corresponding to computing task 3, the video copywriting generation model needs to be preloaded.
[0090] Step S104: Receive the computing task sent by the requester, and the head node submits the computing task to the process corresponding to the computing task for execution.
[0091] When the requester needs to request the execution of a computing task, the requester can issue the computing task. The cluster receives the computing task sent by the requester, and the head node submits the computing task to the process corresponding to the computing task for execution. Since the process corresponding to the computing task has completed the preloading of the required resources, the process can quickly use the preloaded resources to execute the computing task, significantly reducing the startup latency of task execution and effectively improving the efficiency of distributed computing.
[0092] According to the resource preloading method for distributed computing provided by the embodiments of the present application, a cluster including a head node and worker nodes is created. The head node creates a process corresponding to the computing task in the worker nodes according to the process parameters corresponding to the computing task, and downloads the corresponding resources for the process according to the resource download address in the process parameters and preloads the resources, efficiently and conveniently realizing the resource preloading of distributed computing, enabling the process to quickly use the preloaded resources to execute the computing task, significantly reducing the startup latency of task execution, and effectively improving the efficiency of distributed computing.
[0093] Figure 2aA schematic flowchart of a resource preloading method for distributed computing according to another embodiment of the present application is shown, as Figure 2a shown. The method includes the following steps:
[0094] Step S201, create a head node and worker nodes through a cloud platform to form a cluster.
[0095] Specifically, the cloud platform can be a cloud computing platform based on Kubernetes. The head node and worker nodes are created in a containerized environment through the cloud platform, and the cluster is formed by using the head node and worker nodes. In a specific application, the cluster can include one head node and multiple worker nodes, and multiple nodes are managed in the form of a cluster.
[0096] Step S202, set process parameters corresponding to the computing task according to the computing task.
[0097] In the distributed computing framework, a process can be used to execute the computing task. To facilitate resource preloading and computing task execution, the process parameters corresponding to the computing task need to be set according to the task situation of the computing task, etc. Moreover, a task identifier can be generated for the computing task. The task identifier can be a task ID, etc., which is used to uniquely identify the computing task, and different computing tasks have different task identifiers. Considering that there may also be a situation of task version update for the computing task, for example, the task version update caused by the update of the resources required to execute the computing task, in the embodiment of the present application, different task identifiers will be generated for the computing tasks of different task versions, so that the computing tasks can be clearly and conveniently distinguished according to the task identifiers.
[0098] Among them, the process parameters can include the number of replicas, hardware resource quota, runtime environment variables, working directory, dependent data, resource download address, target hash value of the resource, preloaded resource class path, resource class name, etc.
[0099] Specifically, the number of replicas refers to the total number of processes used to execute a computing task. Setting the number of replicas is a strategy to ensure high availability, improve fault tolerance, and enhance performance. More replicas can distribute the request load, improving the service response speed and throughput. The hardware resource quota refers to the configuration of the hardware resources required to start a process, which can specifically include CPU resource quota, GPU resource quota, maximum memory capacity, etc. The dependent data refers to the relevant data of the dependent files required to execute a computing task, which can specifically include the file identifier of the dependent file, the download address of the dependent file, etc. The resource download address refers to the address used to download the resources required during the execution of a computing task. The target hash value of a resource refers to the hash value obtained by performing a hash operation on the resource using a preset hash algorithm. The target hash value has a corresponding relationship with the resource. In a distributed computing framework, resources can be managed in the form of a class. The preloaded resource class path refers to the saving path of the downloaded resources in the worker node. The resource class name refers to the name of the class when the resource is managed in the form of a class.
[0100] In the embodiment of the present application, the singleton class design pattern is adopted to manage resources, which can ensure that there is only one instance of the resource class throughout the program running process and provide a global access point to obtain this instance. Through this processing method, the resources can be conveniently kept resident.
[0101] Step S203: The head node creates a process corresponding to the computing task in the worker node according to the process parameters corresponding to the computing task.
[0102] The head node in the cluster creates a process corresponding to the computing task in the worker node according to the hardware resource quota, runtime environment variables, working directory, and dependent data in the process parameters, realizing the injection of environment variables and configuration when creating the process. Moreover, the head node can also pass the resource download address, the target hash value of the resource, the preloaded resource class path, and the resource class name in the process parameters as parameters to the initialization function corresponding to the process. The initialization function can specifically be the init function. In a programming language, the init function is usually used to perform initialization processing on objects, modules, or classes.
[0103] Step S204: According to the resource download address in the process parameters, the resources required by the process during the execution of the computing task are downloaded to the worker node where the process is located, and the resources are verified.
[0104] Among them, through the initialization function corresponding to the process, according to the resource download address in the process parameters, the resources required by the process during the execution of the computing task can be downloaded to the pre-loaded resource class path of the working node where the process is located. Although the resources are downloaded according to the resource download address, considering that there may be situations such as partial data loss of the downloaded resources and changes in the resources corresponding to the resource download address, in order to ensure that the downloaded resources are the resources required during the execution of the computing task, the resources can be verified again.
[0105] After downloading the resources according to the resource download address, the verification can be carried out in combination with the hash value of the resources. Specifically, using a preset hash algorithm, calculate the current hash value of the resources. For example, perform a hash operation on the downloaded resources using the preset hash algorithm to obtain the current hash value of the resources. Those skilled in the art can select the preset hash algorithm according to actual needs. For example, the preset hash algorithm can be the MD5 algorithm, etc.; after calculating the current hash value of the resources, compare whether the current hash value is consistent with the target hash value of the resources in the process parameters; if the current hash value is consistent with the target hash value, it indicates that the resources obtained through downloading are the resources required during the execution of the computing task, and it is determined that the verification passes; if the current hash value is inconsistent with the target hash value, it indicates that the resources obtained through downloading are different from the resources required during the execution of the computing task, and it is determined that the verification fails.
[0106] Step S205, if the verification passes, it is determined that the process is successfully created, and the singleton class corresponding to the resources is instantiated according to the pre-loaded resource class path and resource class name in the process parameters, and the resources are pre-loaded.
[0107] In the case where the verification of the resources passes, it indicates that the process is successfully created, and the resources required by the process during the execution of the computing task can be pre-loaded. Specifically, through the initialization function corresponding to the process, the path variable in the environment variable can be set according to the pre-loaded resource class path in the process parameters, and the singleton class corresponding to the resources is instantiated according to the resource class name, and the resources are pre-loaded. Among them, the path variable is an environment variable used to specify the search path when the interpreter imports modules.
[0108] Through the above processing method, not only the resource pre-loading is realized, but also the resource class implements the singleton pattern. The singleton class corresponding to the resources is instantiated when the process is created, ensuring that there is only one instance of the singleton class corresponding to the resources throughout the program running process, and providing a global access point to obtain this instance, which conveniently keeps the resources resident. Pre-loading the resources can also reduce the pressure of network transmission and improve the throughput and response speed of the entire system.
[0109] Step S206, add the process reference corresponding to the process to the process reference queue corresponding to the computing task, and manage the process reference queue through the head node.
[0110] After the process is successfully created, it is also necessary to add the process reference corresponding to the process to the process reference queue corresponding to the computing task. Through the process reference, interaction with the process can be carried out, its methods can be called, and results can be obtained. The process reference queue is maintained and managed by the head node in the cluster.
[0111] In the embodiment of the present application, the process reference queue is used to manage the process references of all processes corresponding to the computing task. The process reference queue corresponds to the task identifier of the computing task, so that when it is necessary to allocate the computing task to the process for execution, the head node can conveniently and quickly determine the corresponding process reference queue according to the task identifier of the computing task, and the computing task is submitted to the corresponding process through the process reference in the process reference queue, and the process executes the computing task.
[0112] Step S207, if the verification fails, return a process creation failure notification.
[0113] In the case where the resource verification fails, a process creation failure notification can be returned to inform the head node that the process creation has failed. Subsequently, problems existing in the process creation process can be located and solved through inspection and other means.
[0114] The head node in the cluster creates processes with the number of replicas in the process parameter according to the logic of the above steps S203 to S207, and adds the process references corresponding to each process to the process reference queue corresponding to the computing task, and uses the process reference queue to achieve unified management of all processes corresponding to the computing task.
[0115] Among them, in the process of the head node creating processes with the number of replicas, the head node can select appropriate working nodes to create processes according to the hardware resource usage of each working node in the cluster. Among them, each process corresponding to the computing task can be created in different working nodes in the cluster, and multiple processes can also be created in the same working node. Those skilled in the art can configure the process creation strategy according to actual needs, and no specific limitation is made here. By creating processes in different working nodes, the disaster tolerance ability can be effectively improved. When a certain working node fails, the processes in other working nodes can still continue to provide services to ensure the continuity of the business.
[0116] Figure 2b Shows an architecture schematic diagram of a resource preloading method for distributed computing, as Figure 2bAs shown, the created cluster includes a head node and multiple worker nodes. The head node can, through the interface service, create processes with the number of replicas in multiple worker nodes according to the process parameters corresponding to the computing task, download the resources required by the processes during the execution of the computing task to the worker nodes where the processes are located according to the resource download address in the process parameters, and preload the resources; moreover, the head node is also used to maintain and manage the process reference queue, as Figure 2b shown, the head node maintains a process reference queue corresponding to task identifier 1 and a process reference queue corresponding to task identifier 2. The computing task corresponding to task identifier 1 is different from the computing task corresponding to task identifier 2. For example, the computing task corresponding to task identifier 1 is to remove the watermark in the to-be-processed picture using a watermark removal intelligent model, and the computing task corresponding to task identifier 2 is to add an animation effect to the to-be-processed video.
[0117] In an alternative embodiment, when the head node creates multiple processes in the same worker node of the cluster for the same computing task, it is necessary to download the resources required by each of the multiple processes during the execution of the computing task according to the resource download address in the process parameters. That is to say, the resources used by the multiple processes are independent and there is no sharing of resources, ensuring the independence between the multiple processes.
[0118] In another alternative embodiment, when the head node creates multiple processes in the same worker node of the cluster for the same computing task, it can download the resources required by one of the multiple processes during the execution of the computing task according to the resource download address in the process parameters, and the other processes corresponding to the computing task in this worker node can share the same resource, realizing the sharing of resources by multiple processes corresponding to the same computing task in the same worker node, effectively reducing the storage pressure of the worker node and avoiding unnecessary occupation of storage space on the worker node.
[0119] In the prior art, the life cycle of an ordinary process under a distributed computing framework is usually bound to the process client that creates it. If the process client shuts down or all references to the process go out of scope, the process will be destroyed, thus there are problems of frequent process creation and frequent process destruction. To solve the problems of frequent process creation and frequent process destruction, in the embodiments of the present application, the life cycle of the process corresponding to the computing task is also managed. Specifically, when creating a process, its life cycle attribute can be set to a detached state (such as detached), so as to take over its life cycle and keep the process resident. Among them, the detached state means that the process can exist independently outside the session or process client that creates it, and is not affected by the session or process client that initially creates it, so that the process can still continue to run and is not destroyed after disconnecting from the session or process client that creates it.
[0120] Step S208: Receive the computing task sent by the requestor. The head node selects a target process reference from the process reference queue corresponding to the computing task, and submits the computing task to the process corresponding to the target process reference in the worker node for execution through the target process reference.
[0121] When the requestor needs to request the execution of a computing task, the requestor can issue the computing task according to the task identifier of the computing task. After the cluster receives the computing task sent by the requestor, the head node selects a target process reference from the process reference queue corresponding to the task identifier. Among them, the appropriate process reference corresponding to the process can be selected from the process reference queue according to the task execution progress (such as busy or idle degree, etc.) of each process, the usage of hardware resources of the worker node (such as CPU, GPU, memory, etc.), etc. as the target process reference; through the target process reference, the computing task is submitted to the process corresponding to the target process reference in the worker node for execution.
[0122] As Figure 2b shown, the requestor issues a computing task with a task identifier of 1. The head node selects a target process reference from the process reference queue corresponding to the task identifier 1, and submits the computing task to the process corresponding to the target process reference in the worker node for execution through the interface service.
[0123] Considering that the resources required during the execution of the computing task are still subject to updates, when the resources are updated, the process parameters corresponding to the computing task can be updated so that the resource download address in the process parameters is the download address of the updated resources, and the target hash value is the hash value calculated based on the updated resources. Since the update of the resources also brings about an update of the task version, a new task identifier can be regenerated for the computing task after the task version is updated to effectively distinguish it from the original version. In this case, the head node needs to create a new process corresponding to the computing task and destroy the original process corresponding to the computing task.
[0124] Specifically, when there is an update to the resource download address in the process parameters, the head node creates a new process corresponding to the computing task in the worker nodes according to the process parameters; downloads the corresponding updated resources for the new process according to the updated resource download address in the process parameters, and preloads the updated resources; and the head node destroys the original process corresponding to the computing task.
[0125] The head node in the cluster creates a new process according to the logic of step S203 to step S207 above, and downloads, verifies, and preloads the updated resources, which will not be elaborated here. The head node can conveniently destroy all the processes in the process reference queue corresponding to the original task identifier according to the original task identifier. When the requesting party issues a computing task, the computing task can be issued according to the new task identifier.
[0126] In addition, the embodiment of the present application also provides the ability to dynamically adjust the number of worker nodes in the cluster. The number of worker nodes in the cluster is dynamically adjusted according to the usage of the cluster's hardware resources. Among them, the usage of the cluster's hardware resources can include the usage of hardware resources such as CPU, GPU, and memory of each worker node in the cluster. When the usage of the cluster's hardware resources indicates that the hardware resources of the worker nodes in the cluster are insufficient, the number of worker nodes in the cluster can be increased through the cloud platform; when the usage of the cluster's hardware resources indicates that the idle rate of the hardware resources of the worker nodes in the cluster exceeds the preset threshold, the number of worker nodes in the cluster can be reduced through the cloud platform. Among them, for the worker nodes that need to be destroyed, it is necessary to wait until all the computing tasks running in the worker nodes are completed before destroying them, so as to ensure that the reduction of the number of worker nodes will not affect the normal execution of the computing tasks.
[0127] According to the resource preloading method for distributed computing provided by the embodiments of the present application, the head node in the cluster creates a process corresponding to the computing task in the worker nodes according to the process parameters corresponding to the computing task, downloads the corresponding resources for the process according to the resource download address in the process parameters, and verifies the downloaded resources in combination with the hash value of the resources, effectively ensuring that the downloaded resources are the resources required when executing the computing task, and helping to avoid the situation where the execution of the computing task goes wrong due to incorrect resource download; this solution not only realizes resource preloading, enables the process to quickly utilize the preloaded resources to execute the computing task, significantly reduces the startup latency of task execution, and effectively improves the efficiency of distributed computing, but also the resource class implements the singleton pattern, instantiates the singleton class corresponding to the resource when the process is created, ensures that there is only one instance of the singleton class corresponding to the resource during the entire program operation, and provides a global access point to obtain this instance, conveniently keeping the resources resident; moreover, this solution can also conveniently handle the situation where the resources are updated. When the resources are updated, the head node creates a new process corresponding to the computing task and destroys the original process corresponding to the computing task, realizing the effective management of the process; in addition, this solution also provides the ability to dynamically adjust the number of worker nodes in the cluster, and can dynamically adjust the number of worker nodes in the cluster according to the usage of the cluster hardware resources.
[0128] Figure 3 The structural block diagram of the resource preloading device for distributed computing according to an embodiment of the present application is shown, as Figure 3 shown, the device includes: a cluster creation module 310, a process processing module 320, a preloading module 330, and a task execution module 340.
[0129] The cluster creation module 310 is adapted to: create a cluster including a head node and worker nodes.
[0130] The process processing module 320 is adapted to: the head node creates a process corresponding to the computing task in the worker nodes according to the process parameters corresponding to the computing task set in advance.
[0131] The preloading module 330 is adapted to: download the corresponding resources for the process according to the resource download address in the process parameters and preload the resources.
[0132] The task execution module 340 is adapted to: receive the computing task sent by the requester, and the head node submits the computing task to the process corresponding to the computing task for execution.
[0133] Optionally, the cluster creation module 310 is further adapted to: create the head node and worker nodes through a cloud platform to form a cluster.
[0134] Optionally, the preloading module 330 is further adapted to: download the resources required by the process during the execution of the computing task to the working node where the process is located according to the resource download address in the process parameters, and verify the resources; if the verification passes, determine that the process is successfully created, instantiate the singleton class corresponding to the resources according to the preloading resource class path and resource class name in the process parameters, and preload the resources; if the verification fails, return a process creation failure notification.
[0135] Optionally, the preloading module 330 is further adapted to: calculate the current hash value of the resource using a preset hash algorithm; compare whether the current hash value is consistent with the target hash value of the resource in the process parameters; if the current hash value is consistent with the target hash value, determine that the verification passes; if the current hash value is inconsistent with the target hash value, determine that the verification fails.
[0136] Optionally, the process processing module 320 is further adapted to: after the process is successfully created, add the process reference corresponding to the process to the process reference queue corresponding to the computing task, and manage the process reference queue through the head node.
[0137] Optionally, the task execution module 340 is further adapted to: the head node selects a target process reference from the process reference queue corresponding to the computing task, and submits the computing task to the process corresponding to the target process reference in the working node for execution through the target process reference.
[0138] Optionally, the process processing module 320 is further adapted to: manage the life cycle of the process corresponding to the computing task.
[0139] Optionally, the process processing module 320 is further adapted to: when there is an update to the resource download address in the process parameters, the head node creates a new process corresponding to the computing task in the working node according to the process parameters; the preloading module 330 is further adapted to: download the corresponding updated resources for the new process according to the updated resource download address in the process parameters, and preload the updated resources; the process processing module 320 is further adapted to: the head node destroys the original process corresponding to the computing task.
[0140] Optionally, the device is further adapted to: adjust the number of working nodes in the cluster according to the cluster hardware resource usage of the cluster.
[0141] The descriptions of the above modules refer to the corresponding descriptions in the method embodiments and will not be elaborated here.
[0142] According to the resource preloading device for distributed computing provided by the embodiments of the present application, the head node in the cluster creates a process corresponding to the computing task in the worker nodes according to the process parameters corresponding to the computing task, downloads the corresponding resources for the process according to the resource download address in the process parameters, and verifies the downloaded resources in combination with the hash value of the resources, effectively ensuring that the downloaded resources are the resources required when executing the computing task, and helping to avoid the situation that the execution of the computing task goes wrong due to incorrect resource download; this solution not only realizes resource preloading, enables the process to quickly use the preloaded resources to execute the computing task, significantly reduces the startup latency of task execution, and effectively improves the efficiency of distributed computing, but also the resource class implements the singleton pattern, instantiates the singleton class corresponding to the resource when the process is created, ensures that there is only one instance of the singleton class corresponding to the resource during the entire program running process, and provides a global access point to obtain this instance, conveniently keeping the resources resident; and, this solution can also conveniently handle the situation where the resources are updated. When the resources are updated, the head node creates a new process corresponding to the computing task and destroys the original process corresponding to the computing task, realizing effective management of the process; in addition, this solution also provides the ability to dynamically adjust the number of worker nodes in the cluster, and can dynamically adjust the number of worker nodes in the cluster according to the usage of the cluster hardware resources.
[0143] The embodiments of the present application provide a non-volatile computer storage medium, and the computer storage medium stores at least one executable instruction or computer program, and the executable instruction or computer program can enable the processor to execute the operations corresponding to the resource preloading method for distributed computing in any of the above method embodiments.
[0144] The embodiments of the present application provide a computer program product, and the computer program product includes at least one executable instruction or computer program, and the executable instruction or computer program can enable the processor to execute the operations corresponding to the resource preloading method for distributed computing in any of the above method embodiments.
[0145] Figure 4 The structural schematic diagram of a computing device according to an embodiment of the present application is shown. The specific implementation of the computing device is not limited in the specific embodiments of the present application.
[0146] As Figure 4 shown, the computing device may include: a processor 402, a communications interface 404, a memory 406, and a communication bus 408.
[0147] Among them: The processor 402, the communication interface 404, and the memory 406 communicate with each other through the communication bus 408. The communication interface 404 is used to communicate with network elements of other devices such as clients or other servers. The processor 402 is used to execute the program 410, and specifically can execute the relevant steps in the above-described method embodiments for resource preloading of distributed computing of the computing device.
[0148] Specifically, the program 410 may include program code, and the program code includes computer operation instructions.
[0149] The processor 402 may be a central processing unit CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application. One or more processors included in the computing device may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.
[0150] The memory 406 is used to store the program 410. The memory 406 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk memory.
[0151] The program 410 is specifically used to cause the processor 402 to execute the resource preloading method for distributed computing in any of the above method embodiments. For the specific implementation of each step in the program 410, reference may be made to the corresponding steps and units in the above-described resource preloading embodiments for distributed computing, which will not be elaborated here. Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described devices and modules can refer to the corresponding process descriptions in the foregoing method embodiments, which will not be elaborated here.
[0152] The algorithms and displays provided herein are not inherently related to any specific computer, virtual system, or other device. Various general-purpose systems can also be used in conjunction with the teachings herein. Based on the above description, the structure required to construct such a system is obvious. In addition, the embodiments of the present application are not directed to any specific programming language. It should be understood that the content of the embodiments of the present application described herein can be implemented using various programming languages, and the description of the specific language above is to disclose the best implementation mode of the embodiments of the present application.
[0153] In the specification provided herein, a large number of specific details are set forth. However, it will be understood that embodiments of the present application may be practiced without these specific details. In some instances, well-known methods, structures and techniques have not been shown in detail so as not to obscure the understanding of this specification.
[0154] Similarly, it should be understood that, in order to streamline this disclosure and assist in understanding one or more of the various inventive aspects, in the foregoing description of exemplary embodiments of the present application, the various features of the embodiments of the present application are sometimes grouped together in a single embodiment, figure, or description thereof. However, the disclosed method should not be construed as reflecting an intention that the claimed embodiments of the present application require more features than are expressly recited in each claim. Rather, as reflected in the following claims, the inventive aspects lie in less than all the features of the single foregoing embodiment. Thus, the claims following the detailed description are hereby expressly incorporated into this detailed description, with each claim standing on its own as a separate embodiment of the present application.
[0155] Those skilled in the art will appreciate that the modules in the devices in the embodiments can be adaptively changed and disposed in one or more devices different from those of the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and in addition, they can be divided into multiple sub-modules or sub-units or sub-components. Except that at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all the features disclosed in this specification (including the accompanying claims, abstract and drawings) and all the processes or units of any method or device so disclosed. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstract and drawings) can be replaced by an alternative feature providing the same, equivalent or similar purpose.
[0156] In addition, those skilled in the art will be able to understand that, although some of the embodiments described herein include certain features included in other embodiments but not other features, the combination of the features of different embodiments means that it is within the scope of the embodiments of the present application and forms different embodiments. For example, in the following claims, any one of the claimed embodiments can be used in any combination.
[0157] Each component embodiment of the embodiments of the present application may be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that a microprocessor or a digital signal processor (DSP) may be used in practice to implement some or all of the functions of some or all of the components according to the embodiments of the present application. The embodiments of the present application may also be implemented as a device or apparatus program (such as a computer program and a computer program product) for executing part or all of the methods described herein. Such a program implementing the embodiments of the present application may be stored on a computer-readable medium, or may be in the form of one or more signals. Such signals may be downloaded from an Internet website, or provided on a carrier signal, or in any other form.
[0158] It should be noted that the above embodiments illustrate the embodiments of the present application rather than limit the embodiments of the present application, and those skilled in the art can design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The embodiments of the present application may be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In the unit claims listing several devices, several of these devices may be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words may be interpreted as names.
Claims
1. A resource preloading method for distributed computing, comprising: Creating a cluster including a head node and worker nodes; The head node creates a process corresponding to the computing task in the worker nodes according to the process parameters preset for the computing task; Downloading the corresponding resources for the process according to the resource download address in the process parameters, and preloading the resources; Receiving a computing task sent by a requester, and the head node submitting the computing task to the process corresponding to the computing task for execution.
2. The method according to claim 1, wherein the creating a cluster including a head node and worker nodes further comprises: Creating the head node and the worker nodes through a cloud platform to form the cluster.
3. The method according to claim 1, wherein the downloading the corresponding resources for the process according to the resource download address in the process parameters, and preloading the resources further comprises: According to the resource download address in the process parameters, downloading the resources required by the process during the execution of the computing task to the worker node where the process is located, and verifying the resources; If the verification passes, it is determined that the process is successfully created, and the singleton class corresponding to the resource is instantiated according to the preloading resource class path and resource class name in the process parameters, and the resources are preloaded; If the verification fails, a process creation failure notification is returned.
4. The method according to claim 3, wherein the verifying the resources further comprises: Calculating the current hash value of the resource using a preset hash algorithm; Comparing whether the current hash value is consistent with the target hash value of the resource in the process parameters; If the current hash value is consistent with the target hash value, it is determined that the verification passes; If the current hash value is inconsistent with the target hash value, it is determined that the verification fails.
5. The method according to any one of claims 1-4, the method further comprises: After the process is successfully created, adding the process reference corresponding to the process to the process reference queue corresponding to the computing task, and the head node manages the process reference queue.
6. The method according to claim 5, wherein the head node submitting the computing task to the process corresponding to the computing task for execution further comprises: The head node selects a target process reference from the process reference queue corresponding to the computing task, and through the target process reference, submits the computing task to the process corresponding to the target process reference in the worker node for execution.
7. The method according to any one of claims 1-6, the method further comprises: Managing the life cycle of the process corresponding to the computing task.
8. The method according to any one of claims 1-7, the method further comprises: When there is an update to the resource download address in the process parameters, the head node creates a new process corresponding to the computing task in the worker node according to the process parameters; Download the corresponding updated resources for the new process according to the updated resource download address in the process parameters, and preload the updated resources; Destroy the original process corresponding to the computing task by the head node.
9. The method according to any one of claims 1-8, wherein the method further comprises: Adjust the number of worker nodes in the cluster according to the usage of the cluster hardware resources of the cluster.
10. A resource preloading device for distributed computing, comprising: A cluster creation module adapted to create a cluster including a head node and worker nodes; A process processing module adapted to create a process corresponding to the computing task in the worker nodes by the head node according to the process parameters preset for the computing task; A preloading module adapted to download the corresponding resources for the process according to the resource download address in the process parameters, and preload the resources; A task execution module adapted to receive a computing task sent by a requestor, and submit the computing task to the process corresponding to the computing task by the head node for execution.
11. A computing device, comprising: A processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface complete communication with each other through the communication bus; The memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform the operations corresponding to the resource preloading method for distributed computing according to any one of claims 1-9.
12. A computer storage medium, wherein at least one executable instruction is stored in the storage medium, and the executable instruction causes the processor to perform the operations corresponding to the resource preloading method for distributed computing according to any one of claims 1-9.
13. A computer program product, comprising at least one executable instruction, and the executable instruction causes the processor to perform the operations corresponding to the resource preloading method for distributed computing according to any one of claims 1-9.