Computer resource processing method and device
By analyzing the dependencies of directed acyclic graphs and level 3 cache management, the problem of memory overflow and execution time increase caused by artificial intelligence model loading and unloading strategies is solved, and efficient utilization of memory and system efficiency is achieved.
Patent Information
- Application Number
- CN202510585215.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-05-08
AI Technical Summary
In the prior art, the loading and unloading strategies of artificial intelligence models lead to memory overflow or increased workflow execution time, which cannot effectively balance memory utilization and system efficiency.
By analyzing the dependencies in the directed acyclic graph, we determine whether the dependent resources of the executed node are dependent on the unexecuted node, and free up resources without dependence. We adopt a three-level cache architecture management model, including GPU video memory, CPU memory and disk.
Effectively avoids memory overflow problems, reduces workflow execution time, and improves system efficiency and response speed.
Smart Images

Figure CN120104347B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of computer technologies, and in particular, to a method and apparatus for processing computer resources, a computer device, a computer-readable storage medium, and a computer program product. Background Art
[0002] In the field of artificial intelligence applications, scenarios where multiple artificial intelligence models (such as deep learning models) work together are usually involved. In such scenarios, the loading and unloading of artificial intelligence models are key operations that directly affect the utilization efficiency of system resources and execution performance. Currently, there are mainly two model management strategies: 1. Persistent retention strategy: Once the model is loaded into the GPU video memory, it is not actively released for possible reuse in subsequent steps; 2. Immediate release strategy: The model is immediately released from the GPU video memory after inference is completed and reloaded each time it is used.
[0003] However, although the first strategy can reduce the time overhead of repeated model loading, it is prone to video memory accumulation, resulting in out-of-memory (OOM) errors; although the second strategy can effectively avoid the video memory overflow problem, frequent model loading and unloading operations will significantly increase the workflow execution time and reduce the overall system efficiency.
[0004] It should be noted that the above content is not necessarily prior art and is not used to limit the patent protection scope of the present application. Summary of the Invention
[0005] Embodiments of the present application provide a method and apparatus for processing computer resources, a computer device, a computer-readable storage medium, and a computer program product to solve or alleviate one or more of the above technical problems.
[0006] One aspect of embodiments of the present application provides a method for processing computer resources, the method including:
[0007] Obtain a directed acyclic graph of a target workflow;
[0008] Determine the dependency relationships of each node in the directed acyclic graph, where the dependency relationships include the dependent resources of each node, and the dependent resources include models;
[0009] During the execution of the target workflow, obtain the dependent resources of the executed nodes in the directed acyclic graph as target dependent resources;
[0010] Based on the dependency relationships, determine whether the target dependent resources are depended on by the unexecuted nodes in the directed acyclic graph;
[0011] Release the target dependent resource when the target dependent resource is not depended on by the unexecuted nodes of the directed acyclic graph.
[0012] Optionally, determining the dependency relationships of the nodes in the directed acyclic graph includes:
[0013] Initialize a dependency mapping for explaining each of the dependent resources and the consuming nodes corresponding to the dependent resources;
[0014] Traverse each node in the directed acyclic graph, and during the traversal, determine the dependent resources of the current node as the first dependent resources based on the input of the current node;
[0015] When the first dependent resource is not in the dependency mapping, construct an output key of the first dependent resource and a set of consuming nodes corresponding to the output key in the dependency mapping, and add the current node to the set of consuming nodes, where the output key includes a source node and an output index;
[0016] When the first dependent resource is already in the dependency mapping, add the current node to the set of consuming nodes corresponding to the first dependent resource.
[0017] Optionally, during the execution of the target workflow, obtaining the dependent resources of the executed nodes in the directed acyclic graph as target dependent resources, and determining whether the target dependent resources are depended on by the unexecuted nodes of the directed acyclic graph based on the dependency relationships includes:
[0018] When a first node is executed, traverse the set of consuming nodes corresponding to each output key in the dependency mapping, where the first node is any node in the directed acyclic graph;
[0019] When the first node is included in the currently traversed set of consuming nodes, delete the first node from the currently traversed set of consuming nodes;
[0020] When the currently traversed set of consuming nodes is empty after deletion, determine that the dependent resources corresponding to the currently traversed set of consuming nodes are not depended on by the unexecuted nodes of the directed acyclic graph.
[0021] Optionally, releasing the target dependent resource when the target dependent resource is not depended on by the unexecuted nodes of the directed acyclic graph includes:
[0022] When the target dependent resource is not depended on by the unexecuted nodes of the directed acyclic graph, determine whether the video memory pressure of the current GPU video memory is less than a first pressure threshold;
[0023] When the video memory pressure is less than the first pressure threshold, skip the step of releasing the target dependent resource.
[0024] Optionally, the method further includes:
[0025] When the video memory pressure is greater than or equal to the first pressure threshold, determine whether the video memory pressure is less than the second pressure threshold;
[0026] When the video memory pressure is less than the second pressure threshold, transfer the target dependent resource from the GPU video memory to the CPU memory;
[0027] When the video memory pressure is greater than or equal to the second pressure threshold, transfer the target dependent resource from the GPU video memory to the disk.
[0028] Optionally, the method further includes:
[0029] Before executing the second node of the directed acyclic graph, determine the dependent resource corresponding to the second node as the second dependent resource, where the second node is any node of the directed acyclic graph;
[0030] Check in sequence whether the second dependent resource is in the GPU video memory, the CPU memory or the disk;
[0031] When the second dependent resource is in the CPU memory or the disk, load the second dependent resource into the GPU video memory.
[0032] Optionally, the method further includes:
[0033] Intercept the target function from a third-party node, where the target function includes a model loading function, a model unloading function and a migration operation function;
[0034] When the target function is intercepted, add the workflow corresponding to the target function to the directed acyclic graph of the target workflow.
[0035] Another aspect of the embodiments of the present application provides a computer resource processing device, where the device includes:
[0036] A first acquisition module, configured to acquire a directed acyclic graph of a target workflow;
[0037] A first determination module, configured to determine the dependency relationship of each node in the directed acyclic graph, where the dependency relationship includes the dependent resources of each node, and the dependent resources include models;
[0038] A second acquisition module, configured to, during the execution of the target workflow, acquire the dependent resources of the executed nodes in the directed acyclic graph as target dependent resources;
[0039] A second determination module, configured to determine, based on the dependency relationship, whether the target dependent resource is depended on by unexecuted nodes of the directed acyclic graph;
[0040] A processing module, configured to release the target dependent resource when the target dependent resource is not depended on by unexecuted nodes of the directed acyclic graph.
[0041] Another aspect of the embodiments of the present application provides a computer device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein: the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method as described above.
[0042] Another aspect of the embodiments of the present application provides a computer-readable storage medium, in which computer instructions are stored, and when the computer instructions are executed by a processor, the method as described above is implemented.
[0043] Another aspect of the embodiments of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method as described above is implemented.
[0044] The embodiments of the present application adopting the above technical solutions may include the following advantages:
[0045] By obtaining the directed acyclic graph of the target workflow, determining the dependency relationships of the nodes in the directed acyclic graph, where the dependency relationships include the dependent resources of each node, and the dependent resources include models; during the execution of the target workflow, obtaining the dependent resources of the executed nodes in the directed acyclic graph as the target dependent resources, determining whether the target dependent resources are depended on by unexecuted nodes of the directed acyclic graph based on the dependency relationships, and releasing the target dependent resources when the target dependent resources are not depended on by unexecuted nodes of the directed acyclic graph, it is possible to analyze the dependency relationships of the nodes in the workflow directed acyclic graph, determine whether the resources (including models) in the system are depended on by unexecuted nodes according to the dependency relationships, and release the resources when not depended on, which can effectively avoid problems such as video memory overflow in scenarios such as multi-model collaboration, and at the same time can reasonably reduce the workflow execution time, improve the response speed of requests, and improve system efficiency. Description of the Drawings
[0046] The drawings exemplarily show the embodiments and form a part of the specification, and are used together with the written description of the specification to explain the exemplary embodiments of the embodiments. The shown embodiments are for illustrative purposes only and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0047] Figure 1 Schematically shows the operating environment diagram of the computer resource processing method according to Embodiment 1 of the present application;
[0048] Figure 2 Schematically shows the flowchart of the computer resource processing method according to Embodiment 1 of the present application;
[0049] Figure 3 Schematically shows Figure 2 The sub-step flowchart of step S202 in;
[0050] Figure 4 Schematically shows Figure 2 The sub-step flowchart of steps S204 and S206 in;
[0051] Figure 5 Schematically shows Figure 2 The sub-step flowchart of step S208 in;
[0052] Figure 6 Schematically shows the new process of the computer resource processing method according to Embodiment 1 of the present application;
[0053] Figure 7 Schematically shows another new process of the computer resource processing method according to Embodiment 1 of the present application;
[0054] Figure 8 Schematically shows yet another new process of the computer resource processing method according to Embodiment 1 of the present application;
[0055] Figure 9 Schematically shows the system architecture diagram of the computer resource processing method according to Embodiment 1 of the present application;
[0056] Figure 10 Schematically shows the timing process example diagram of the computer resource processing method according to Embodiment 1 of the present application;
[0057] Figure 11 Schematically shows the block diagram of the computer resource processing device according to Embodiment 2 of the present application; and
[0058] Figure 12 Schematically shows the hardware architecture schematic diagram of the computer device according to Embodiment 3 of the present application. Detailed implementation manners
[0059] To make the objectives, technical solutions and advantages of this application more clear and understandable, the following further elaborates on this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts belong to the scope of protection of this application.
[0060] It should be noted that the descriptions involving "first", "second", etc. in the embodiments of this application are only for descriptive purposes and cannot be construed as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. Additionally, the technical solutions between various embodiments can be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions results in contradictions or inability to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by this application.
[0061] In the description of this application, it should be understood that the numerical labels before the steps do not identify the sequence of execution of the steps, but are only used to facilitate the description of this application and to distinguish each step. Therefore, it cannot be construed as a limitation to this application.
[0062] First, the following provides the term explanations related to this application:
[0063] Workflow Engine: A software component or system used to define, manage, and automatically execute workflows. By decomposing business processes into a series of executable task nodes and driving the task flow according to preset rules and logics (such as sequence, conditional branching, parallel processing, etc.), it realizes the automated management of business processes.
[0064] Directed Acyclic Graph (DAG): The basic data structure of a workflow, representing the dependency relationships between nodes, ensuring that nodes are executed in the correct order and do not form circular dependencies.
[0065] Video Memory Overflow: A common technical problem in computer graphics, deep learning, or high-performance computing, referring to an abnormal state where the video memory (Video RAM, VRAM) of a graphics processing unit (GPU) cannot accommodate the currently required data or computing tasks, resulting in the exhaustion of memory resources.
[0066] Community Node: A workflow node created by third-party developers.
[0067] Function Interception: It is a programming technique that refers to inserting custom logic before and after (or during the execution of) a target function to monitor, modify, or control the behavior of the function without directly modifying the source code of the target function. This technique is commonly used in scenarios such as debugging, logging, performance analysis, permission control, input / output filtering, etc. The core goal is to extend or change the program's functionality without invading the original code.
[0068] Dependency Graph: It is a directed graph used to describe the dependency relationships between components, modules, tasks, data, or variables in a system. It is commonly used in fields such as software engineering, project management, compilation principles, data processing, etc. The core idea is to visualize the dependency relationships between elements through a graph structure to help analyze issues such as the complexity of the system, execution order, conflicts, or circular dependencies.
[0069] Hook Function: A callback function triggered by a specific event, used to implement custom logic such as resource management.
[0070] Secondly, to facilitate the understanding of the technical solutions provided in the embodiments of the present application by those skilled in the art, the related technologies are described below:
[0071] In the scenario where multiple artificial intelligence models work collaboratively, the loading and unloading of artificial intelligence models are relatively critical operations, which directly affect the utilization efficiency of system resources and execution performance. The two main current model management strategies are: 1. Persistent Retention Strategy: Once the model is loaded into the GPU video memory, it is not actively released; 2. Immediate Release Strategy: The model is immediately released from the GPU video memory every time inference is completed and reloaded when used next time.
[0072] However, although the first strategy can reduce the time overhead of repeated model loading, it is prone to video memory accumulation and video memory overflow errors; although the second strategy can avoid video memory overflow, it will significantly increase the workflow execution time and reduce system efficiency.
[0073] Therefore, the embodiments of the present application provide a technical solution for computer resource processing. In this technical solution, by analyzing the dependency relationships of each node in the workflow directed acyclic graph, and determining whether the resources (including models) in the system are depended on by unexecuted nodes according to the dependency relationships, and releasing the resources when they are not depended on, it can effectively avoid problems such as video memory overflow in scenarios such as multi-model collaboration, and at the same time can reasonably reduce the workflow execution time and improve system efficiency. See the following for details.
[0074] Finally, for ease of understanding, an exemplary operating environment is provided below.
[0075] Such asFigure 1 As shown, the environmental schematic diagram includes a service platform 2, a network 4, and a client 6, where:
[0076] The service platform 2 can be composed of a single or multiple computing devices. The multiple computing devices can include virtualized computing instances. The virtualized computing instances can include virtual machines, such as emulations of computer systems, operating systems, servers, etc. The computing devices can load virtual machines based on virtual images and / or other data that define specific software (e.g., operating systems, dedicated applications, servers) for emulation. As the demand for different types of processing services changes, different virtual machines can be loaded and / or terminated on one or more computing devices. A hypervisor can be implemented to manage the use of different virtual machines on the same computing device.
[0077] The service platform 2 can be configured to communicate with the client 6, etc. via the network 4. The network 4 includes various network devices, such as routers, switches, multiplexers, hubs, modems, bridges, repeaters, firewalls, proxy devices, and / or the like. The network 4 can include physical links, such as coaxial cable links, twisted pair cable links, fiber optic links and their combinations, etc., or wireless links, such as cellular links, satellite links, Wi-Fi links, etc.
[0078] The service platform 2 can provide services such as storage, reading, writing, querying, deleting, etc., such as providing services like drawing and video production for the client.
[0079] The client 6 can be an electronic device running an operating system such as Windows, Android™, or iOS, such as a smartphone, tablet device, laptop computer, virtual reality device, gaming device, set-top box, in-vehicle terminal, smart TV. Based on the above operating systems, various application programs can be run, such as running application programs like drawing and video production.
[0080] The client 6 can provide / configure a user access page for manipulating the service platform 2 or uploading objects, etc.
[0081] It should be noted that the above devices are exemplary, and the number and types of devices can be adjusted in different scenarios or according to different requirements.
[0082] The technical solutions of the present application will be introduced below through multiple embodiments. It should be noted that these embodiments can be implemented in many different forms and should not be construed as being limited only to the embodiments described herein.
[0083] Embodiment 1
[0084] Figure 2The flowchart of the computer resource processing method according to Embodiment 1 of the present application is schematically shown. It should be noted that the execution subject of the computer resource processing method in the embodiments of the present application can be Figure 1 a service platform, or more specifically, a computing node in the service platform.
[0085] As Figure 2 shown, the computer resource processing method may include steps S200 to S208, where:
[0086] Step S200: Obtain the directed acyclic graph of the target workflow.
[0087] Step S202: Determine the dependency relationships of the nodes in the directed acyclic graph. The dependency relationships include the dependent resources of each node, and the dependent resources include models.
[0088] Step S204: During the execution of the target workflow, obtain the dependent resources of the executed nodes in the directed acyclic graph as the target dependent resources.
[0089] Step S206: Based on the dependency relationships, determine whether the target dependent resources are dependent on the unexecuted nodes of the directed acyclic graph.
[0090] Step S208: When the target dependent resources are not dependent on the unexecuted nodes of the directed acyclic graph, release the target dependent resources.
[0091] The computer resource processing method provided in this embodiment, by obtaining the directed acyclic graph of the target workflow, determines the dependency relationships of the nodes in the directed acyclic graph. The dependency relationships include the dependent resources of each node, and the dependent resources include models; during the execution of the target workflow, obtain the dependent resources of the executed nodes in the directed acyclic graph as the target dependent resources, and based on the dependency relationships, determine whether the target dependent resources are dependent on the unexecuted nodes of the directed acyclic graph. When the target dependent resources are not dependent on the unexecuted nodes of the directed acyclic graph, release the target dependent resources. It can determine whether the resources (including models) in the system are dependent on the unexecuted nodes according to the dependency relationships by analyzing the dependency relationships of the nodes in the workflow directed acyclic graph, and release the resources when they are not dependent, which can effectively avoid problems such as video memory overflow in scenarios such as multi-model collaboration, and at the same time can reasonably reduce the execution time of the workflow, improve the response speed of requests and system efficiency.
[0092] The following will Figure 2 be elaborated in detail on each step and optional other steps in steps S200 to S208.
[0093] Step S200 , obtain the directed acyclic graph of the target workflow.
[0094] Among them, the target workflow can be a specific or arbitrary workflow in the system. The directed acyclic graph of the target workflow can be constructed or parsed and generated according to the specific business logic of the target workflow (such as code, configuration file, visual metadata). In practical applications, the directed acyclic graph of the target workflow can be completed by a workflow engine, and when obtaining the directed acyclic graph of the target workflow, it can be obtained through the workflow engine.
[0095] Step S202 , determine the dependency relationships of each node in the directed acyclic graph. The dependency relationships include the dependent resources of each node, and the dependent resources include models.
[0096] Specifically, the dependent resources of each node can be determined according to the inputs of each node in the directed acyclic graph and the connection relationships of each node. Among them, the dependent resources can be the output of a certain node. For example, if a certain node a is to stylize the image of the previous node by calling the A model, the dependent resources generated by this node are the A model and the image of the previous node, and its output is the stylized image; if the input of a certain node b is this stylized image, the dependent resources of node b can be this stylized image; for another example, if node c is to load the A model, the A model is the output of node c. By performing the same analysis on all nodes in the directed acyclic graph, the dependency relationships of each node can be determined. Among them, the dependency relationships can specifically be represented in forms such as dependency lists, dependency trees, dependency graphs, and dependency networks, and there is no limit here.
[0097] Step S204 , during the execution of the target workflow, obtain the dependent resources of the executed nodes in the directed acyclic graph as the target dependent resources.
[0098] Specifically, it can be after a certain node is executed, obtain its corresponding dependent resources as the target dependent resources to determine whether the target dependent resources are depended on by other unexecuted nodes. For example, if node d has been executed, and the dependent resource of node d is the B model, then the B model can be used as the target dependent resource. Optionally, it can also be after a certain node is executed, obtain the dependent resources of all the executed nodes as the target dependent resources. For example, if the nodes a, b, c, d of the directed acyclic graph have been executed, then the dependent resources corresponding to nodes a, b, c, d can be obtained as the target dependent resources, that is, comprehensively check the dependent resources of all the executed nodes.
[0099] Step S206 , based on the dependency relationships, determine whether the target dependent resources are depended on by the unexecuted nodes of the directed acyclic graph.
[0100] Specifically, it is possible to determine whether the target dependent resource is dependent on unexecuted nodes according to the dependency relationships determined through previous analysis. For example, if it is found through the dependency relationships that the target dependent resource, the B model, is not required by any unexecuted nodes, it is determined that the B model is not dependent on unexecuted nodes; if it is found through the dependency relationships that the B model is dependent on the unexecuted node e, it is determined that the B model is dependent on unexecuted nodes.
[0101] Step S208 , in the case where the target dependent resource is not dependent on unexecuted nodes of the directed acyclic graph, release the target dependent resource.
[0102] Specifically, if it is determined that the target dependent resource is not dependent on unexecuted nodes, the target dependent resource can be released, thereby freeing up space, reducing the accumulation of GPU video memory, and avoiding the problem of video memory overflow. For example, if the target dependent resource, the B model, is not dependent on unexecuted nodes, the B model can be released. Of course, in the case where it is determined that the target dependent resource is not dependent on unexecuted nodes, other conditions can also be judged (such as whether the video memory is sufficient). Only when other conditions are met is the target dependent resource released; otherwise, the target dependent resource is retained. For example, a condition judgment on whether the video memory is sufficient can be made. When the video memory is insufficient, the target dependent resource is released; when the video memory is sufficient, the target dependent resource is retained for possible subsequent use.
[0103] In an alternative embodiment, as Figure 3 shown, in step S202, determining the dependency relationships of each node in the directed acyclic graph may include:
[0104] Step S300, initialize a dependency map, which is used to illustrate each dependent resource and the consuming nodes corresponding to the dependent resources.
[0105] Step S302, traverse each node in the directed acyclic graph, and during the traversal, determine the dependent resource of the current node as the first dependent resource based on the input of the current node.
[0106] Step S304, in the case where the first dependent resource is not in the dependency map, construct an output key of the first dependent resource and a set of consuming nodes corresponding to the output key in the dependency map, and add the current node to the set of consuming nodes. The output key includes a source node and an output index.
[0107] Step S306, in the case where the first dependent resource is already in the dependency map, add the current node to the set of consuming nodes corresponding to the first dependent resource.
[0108] Specifically, a blank dependency map can be initialized, and then traversal starts from the first node of the directed acyclic graph. For each traversed node, its input is analyzed to determine the dependency resource corresponding to the currently traversed node as the first dependency resource. In the case where the first dependency resource is not in the dependency map, an output key of the first dependency resource and a set of consuming nodes corresponding to the output key are constructed in the dependency map, and the current node is added to the currently constructed set of consuming nodes. In the case where the first dependency resource is already in the dependency map, the current node is added to the set of consuming nodes corresponding to the output key of the first dependency resource. When all nodes have been traversed, a complete dependency map can be obtained. For example, if the first dependency resource corresponding to the current node is Model C, and Model C is the first output of Node a, and the first dependency resource is not in the dependency map, then an output key of Model C can be constructed in the dependency map as (Node a, 1), the set of consuming nodes is empty, and then the current node is added to the set of consuming nodes of Model C.
[0109] In this embodiment, by initializing the dependency map and traversing each node of the directed acyclic graph, during the traversal process, the dependency resource is determined based on the input of the current node as the first dependency resource. In the case where the first dependency resource is not in the dependency map, an output key of the first dependency resource and the corresponding set of consuming nodes are constructed in the dependency map, and the current node is added to the set of consuming nodes. In the case where the first dependency resource is already in the dependency map, the current node is added to the corresponding consuming node, which can effectively construct the dependency relationships of each node in the directed acyclic graph and facilitate subsequent judgment on whether to release resources.
[0110] In an alternative embodiment, in steps S204 and S206, that is, during the execution of the target workflow, the dependency resources of the executed nodes in the directed acyclic graph are obtained as the target dependency resources, and based on the dependency relationships, it is determined whether the target dependency resources are depended on by the unexecuted nodes of the directed acyclic graph. As Figure 4 shown, it may include:
[0111] Step S400, when the first node has been executed, traverse the set of consuming nodes corresponding to each output key in the dependency map, where the first node is any node in the directed acyclic graph.
[0112] Step S402, when the first node is included in the currently traversed set of consuming nodes, remove the first node from the currently traversed set of consuming nodes.
[0113] Step S404, when the currently traversed set of consuming nodes is empty after deletion, determine that the dependency resources corresponding to the currently traversed set of consuming nodes are not depended on by the unexecuted nodes of the directed acyclic graph.
[0114] Specifically, when the execution of the first node (any node) is completed, traverse the set of consuming nodes corresponding to each output key in the dependency mapping. If the first node is included in the currently traversed set of consuming nodes, remove the first node from the currently traversed set of consuming nodes. After the removal, judge the currently traversed set of consuming nodes. If the currently traversed set of consuming nodes is empty after the removal, it is determined that the dependency resource corresponding to the currently traversed set of consuming nodes is not depended on by the unexecuted nodes of the directed acyclic graph. If the currently traversed set of consuming nodes is not empty after the removal, it is determined that the dependency resource corresponding to the currently traversed set of consuming nodes is depended on by the unexecuted nodes of the directed acyclic graph. For example, after the execution of the first node is completed, traverse the set of consuming nodes in the dependency mapping. If the first node is included in the set of consuming nodes of model A, remove the first node from the set of consuming nodes of model A. If the set of consuming nodes corresponding to model A is empty after the removal, it is determined that model A is not depended on by the unexecuted nodes. Among them, a hook function can be added after the nodes in the directed acyclic graph, so that after a node is executed, the traversal and subsequent operations can be started through the hook function. In practical applications, in order to facilitate the release and other operations of the resources not depended on by the unexecuted nodes, when it is determined that the dependency resource corresponding to the currently traversed set of consuming nodes is not depended on by the unexecuted nodes, the dependency resource can be marked as releasable or deletable. Optionally, the dependency resources in the dependency mapping can also be reference-counted according to the dependent nodes. After a reference node is executed, the corresponding number of times of the dependency resource is reduced by 1. Finally, it can be determined whether the dependency resource is depended on by the unexecuted nodes according to whether the reference count of the dependency resource is zero.
[0115] In this embodiment, when any node in the directed acyclic graph is executed, traverse the set of consuming nodes in the dependency mapping, remove the executed node from the set of consuming nodes, and then determine whether the corresponding dependency resource is depended on by the unexecuted nodes according to whether the set of consuming nodes after the removal is empty, so as to effectively determine whether the resources in the system still need to run in the GPU video memory, so as to accurately determine the resources that can be released.
[0116] In an alternative embodiment, in step S208, that is, when the target resource is not depended on by the unexecuted nodes of the directed acyclic graph, release the target dependency resource, as Figure 5 shown, it may include:
[0117] Step S500, when the target dependency resource is not depended on by the unexecuted nodes of the directed acyclic graph, determine whether the video memory pressure of the current GPU video memory is less than the first pressure threshold.
[0118] Step S502, when the video memory pressure is less than the first pressure threshold, skip the release step of the target dependency resource.
[0119] The first pressure threshold can be a value used to indicate whether the GPU video memory is sufficient. That is, if the video memory pressure is less than the first pressure threshold, it means the video memory is sufficient; if the video memory pressure is greater than or equal to the first pressure threshold, it means the video memory is insufficient. The setting of the first pressure threshold can be determined according to specific application scenarios, hardware configurations, and strategies for using the GPU video memory, and no specific restrictions are made here. Among them, the video memory pressure can be comprehensively determined based on indicators such as the absolute size of the used video memory and the percentage of the used video memory in the total video memory.
[0120] Specifically, when it is determined that the target dependent resource is not dependent on the unexecuted nodes of the directed acyclic graph, further determine whether the video memory pressure of the current GPU video memory is less than the first pressure threshold; when the video memory pressure is less than the first pressure threshold, it is determined that the current video memory is sufficient, and the release step of the target dependent resource can be skipped, and the target dependent resource can be temporarily retained in the GPU video memory.
[0121] In this embodiment, by determining whether the video memory pressure of the current GPU video memory is less than the first pressure threshold when the target dependent resource is not dependent on the unexecuted nodes of the directed acyclic graph, and skipping the release step of the target dependent resource when the video memory pressure is less than the first pressure threshold, the resource can be temporarily retained in the GPU video memory when the video memory is sufficient, so that it can be directly used without reloading when needed later, reducing the startup and shutdown overhead.
[0122] In an alternative embodiment, as Figure 6 shown, the computer resource processing method of the embodiment of the present application may further include:
[0123] Step S600, when the video memory pressure is greater than or equal to the first pressure threshold, determine whether the video memory pressure is less than the second pressure threshold.
[0124] Step S602, when the video memory pressure is less than the second pressure threshold, transfer the target dependent resource from the GPU video memory to the CPU memory.
[0125] Step S604, when the video memory pressure is greater than or equal to the second pressure threshold, transfer the target dependent resource from the GPU video memory to the disk.
[0126] The second pressure threshold can be a value used to indicate whether the video memory pressure is moderate. That is, if the video memory pressure is less than the second pressure threshold, it means that although the video memory is not sufficient, the video memory pressure is still in a moderate state, which can be understood as medium video memory pressure; if the video memory pressure is greater than or equal to the second threshold, it means that the video memory pressure is high. Similarly, the setting of the second pressure threshold can be determined according to specific application scenarios, hardware configurations, and strategies for using the GPU video memory, and no specific restrictions are made here. Among them, the second pressure threshold is greater than the first pressure threshold.
[0127] Specifically, when the video memory pressure is greater than or equal to the first pressure threshold, it is further determined whether the video memory pressure is less than the second pressure threshold; if the video memory pressure is less than the second pressure threshold, it indicates that the video memory pressure is at a medium level at this time, then the target dependent resource can be transferred from the GPU video memory to the CPU memory, specifically, the target dependent resource can be released from the GPU video memory, and the running location of the target dependent resource can be migrated to the CPU memory; if the video memory pressure is greater than or equal to the second pressure threshold, it indicates that the video memory pressure is high at this time, then the target dependent resource can be transferred from the GPU video memory to the disk, specifically, the target dependent resource can be released from the GPU video memory, and the running location of the target dependent resource can be migrated to the disk.
[0128] In this embodiment, when the video memory pressure is greater than or equal to the first pressure threshold, it is determined whether the video memory pressure is less than the second pressure threshold. When the video memory pressure is less than the second pressure threshold, the target dependent resource is transferred from the GPU video memory to the CPU memory; when the video memory pressure is greater than or equal to the second pressure threshold, the target dependent resource is transferred from the GPU video memory to the disk. Resources can be migrated to the CPU memory when the video memory pressure is medium, and subsequent loading from the CPU memory can accelerate the response speed; at the same time, resources can also be migrated to the disk when the pressure is high to reduce the video memory pressure, realizing multi-level caching under different video memory pressures and improving the refinement degree of resource management.
[0129] In an alternative embodiment, as Figure 7 shown, the computer resource processing method of the embodiment of the present application may further include:
[0130] Step S700, before executing the second node of the directed acyclic graph, determine the dependent resource corresponding to the second node as the second dependent resource, where the second node is any node of the directed acyclic graph.
[0131] Step S702, sequentially check whether the second dependent resource is in the GPU video memory, the CPU memory, or the disk.
[0132] Step S704, when the second dependent resource is in the CPU memory or the disk, load the second dependent resource into the GPU video memory.
[0133] Among them, by adding hook functions in front of all nodes of the directed acyclic graph, the content of step S700 and subsequent steps can be started before executing the nodes.
[0134] Specifically, before executing the second node of the directed acyclic graph, determine the dependent resource corresponding to the second node as the second dependent resource; then, sequentially check whether the second dependent resource is in the GPU video memory, CPU memory, or disk; in the case where the second dependent resource is in the GPU video memory, the second dependent resource can be directly used; in the case where the second dependent resource is in the CPU memory or disk, the second dependent resource is loaded into the GPU video memory for use.
[0135] In this embodiment, before executing any node of the directed acyclic graph, determine the dependent resource corresponding to the current node, sequentially check whether the corresponding dependent resource is in the GPU video memory, CPU memory, or disk, and in the case where the corresponding dependent resource is in the CPU memory or disk, load the corresponding dependent resource into the GPU video memory, so that resources can be loaded from an appropriate cache level, and at the same time, corresponding preparations can be made in advance for the execution of the node, improving the execution efficiency of the node.
[0136] In an alternative embodiment, as Figure 8 shown, the computer resource processing method of the embodiment of the present application may further include:
[0137] Step S800, intercept the target function from a third-party node, where the target function includes model loading, model unloading, and migration operation functions.
[0138] Step S802, in the case of intercepting the target function, add the workflow corresponding to the target function to the directed acyclic graph of the target workflow.
[0139] Specifically, a function interception can be set in the access path of a third-party node (such as a community node). When the functions in the third-party node include functions such as model loading, model unloading, and model migration, determine them as target functions for interception; then parse the workflows corresponding to these target functions, and add the workflows corresponding to the target functions to the directed acyclic graph of the target workflow. Subsequently, the corresponding models can be incorporated into the dependency relationship for management, so as to achieve the purpose of unified resource management of the models corresponding to the third-party nodes. Optionally, in order to precisely manage these resources of the model, model identifiers can be generated according to various factors such as model source, structural characteristics, and initialization parameters. When one factor is different, the generated model identifiers are also different.
[0140] In this embodiment, by intercepting target functions such as model loading, model unloading, and migration operations from a third-party node and adding the intercepted target functions to the directed acyclic graph of the target workflow, the corresponding model management can be incorporated into unified resource management without modifying the target functions of the third-party node, thereby achieving resource optimization to a greater extent.
[0141] To make this application easier to understand, the following provides an exemplary application in conjunction with Figure 9 and Figure 10 . Among them, Figure 9 is the system architecture diagram corresponding to the computer resource processing method of the embodiment of this application, Figure 10 is the timing process example diagram of the computer resource processing method.
[0142] I. As shown in Figure 9 , the system mainly includes a workflow engine executor, a model manager, and an engine optimizer.
[0143] 1. Workflow engine executor: Responsible for parsing and executing the DAG workflow. This module receives the user-defined workflow graph and executes the computing tasks of each node according to the dependency relationship between nodes.
[0144] Specifically, the workflow engine executor can implement the following functions:
[0145] ⑴ Workflow DAG topological sorting to determine the node execution order;
[0146] ⑵ Data transfer and parameter mapping between nodes;
[0147] ⑶ Node execution status management and exception handling.
[0148] 2. Model manager: Responsible for model loading, unloading, and cache management, and implements a three-level cache architecture, namely:
[0149] ⑴ L0 cache: GPU video memory, used for directly executing model inference, with the fastest access speed;
[0150] ⑵ L1 cache: CPU memory, used for temporarily storing models that are not currently in use but may be called soon;
[0151] ⑶ L2 cache: Disk storage, which stores the original weight files of the models.
[0152] The model manager provides unified model loading and unloading interfaces for the adapted standard nodes, and at the same time intercepts the model operations of the community nodes that have not been fully adapted through the hook mechanism to incorporate them into unified management.
[0153] Key mechanisms of the model manager:
[0154] ⑴ Model loading optimization: Adopt the optimal loading path according to the position of the model in the three-level cache;
[0155] ⑵ Model unloading strategy: Perform precise unloading based on the dependency analysis results;
[0156] ⑶ Cache level conversion: Dynamically adjust the storage position of the model among the three-level caches according to the workflow execution status and video memory pressure;
[0157] ⑷ Model status tracking: Real-time monitoring of the model's loading status, usage status, and cache location.
[0158] In implementation, the model manager loads the model from an appropriate cache level according to the requirements of the workflow nodes. First, it checks the L0 cache. If it does not exist, it then checks the L1 and L2 caches in turn. This mechanism effectively reduces the overhead of repeated model loading.
[0159] 3. Engine optimizer: It is a bridge connecting the workflow engine executor and the model manager, responsible for implementing intelligent resource management of the workflow. This module achieves resource optimization through three key mechanisms:
[0160] ⑴ Non-standard node interception: Intercept the model loading and unloading logic of community nodes and redirect it to the model manager for unified management. The interception mechanism does not modify the original node code, but dynamically takes over resource management operations at runtime through function hook technology.
[0161] ⑵ Workflow dependency analysis: Before the workflow is executed, analyze and construct a dependency graph between nodes, and record which subsequent nodes reference the output of each node. This information is the basis for performing dynamic pruning (resource release).
[0162] ⑶ Node execution hook: Implement hook functions before and after each node of the workflow for resource preloading and automatic pruning. After the node is executed, the hook checks the dependency graph, identifies resources that are no longer needed, and triggers the unloading operation.
[0163] The dependency analysis and dynamic pruning in the embodiments of this application are the core of achieving resource optimization, mainly including the following links:
[0164] ⑴ Static dependency analysis: Before the workflow is executed, perform a depth-first traversal of the DAG to identify the input-output dependency relationships between nodes. For the output of each node, record all subsequent nodes that depend on this output to construct a complete dependency relationship (dependency graph).
[0165] ⑵ Dynamic execution tracking: During the execution of the workflow, continuously update the node execution status and mark the completed nodes. Combining the dependency graph information, calculate the active status of each resource in real time.
[0166] ⑶ Intelligent pruning decision: After the node is executed, check the dependency situation of its input and output. If an input or output is no longer referenced by unexecuted nodes, mark the resources related to this input or output as releasable. According to the resource characteristics and system status, decide whether to perform the actual release operation.
[0167] ⑷ Adaptive pruning strategy: When system resources are sufficient, pruning is skipped to increase the probability that the model cache will be utilized by the next workflow; when resources are scarce, more aggressive pruning is performed.
[0168] II. As Figure 10 shown, the computer resource processing method generally includes the following processes:
[0169] 1. The user submits a workflow DAG;
[0170] 2. The workflow engine requests dependency analysis from the engine prioritizer, and the engine optimizer performs dependency analysis to construct a dependency graph;
[0171] 3. The engine prioritizer returns the constructed dependency graph to the workflow engine;
[0172] 4. For each node to be executed in the DAG, the workflow engine sends a request to the engine optimizer through the pre-execution hook, and the engine optimizer notifies the model manager to prepare model resources;
[0173] 5. The model manager first determines whether the required model is in the L0 cache. If it is, it is directly used; if it is in the L1 cache, it is loaded into the GPU; if it is in the L2 cache, the model is read and loaded into the GPU;
[0174] 6. The resources are ready and the pre-processing is completed.
[0175] 7. After the node is executed, the dependency status is updated through the post-execution hook via the engine optimizer, pruning decision-making is performed, and the model manager is requested to release the specified resources;
[0176] 8. The model manager determines whether the current video memory is sufficient. If it is sufficient, the release of resources is skipped and the resources are retained in the L0 cache; if the video memory pressure is medium, the resources are transferred to the L1 cache; if the video memory pressure is high, the resources are transferred to the L2 cache;
[0177] 9. The release of resources is completed;
[0178] 10. The post-processing is completed;
[0179] 11. After the workflow is executed, the workflow engine returns the workflow result to the user.
[0180] Embodiment II
[0181] Figure 11Schematically shown is a block diagram of a computer resource processing apparatus according to Embodiment 2 of the present application. The apparatus can be divided into one or more program modules. One or more program modules are stored in a storage medium and executed by one or more processors to complete the embodiments of the present application. The program modules referred to in the embodiments of the present application refer to a series of computer program instruction segments that can complete specific functions. The following description will specifically introduce the functions of each program module in this embodiment. As Figure 11 shown, the apparatus 900 may include: a first acquisition module 910, a first determination module 920, a second acquisition module 930, a second determination module 940, and a processing module 950, where:
[0182] The first acquisition module 910 is configured to acquire a directed acyclic graph of a target workflow;
[0183] The first determination module 920 is configured to determine the dependency relationships of each node in the directed acyclic graph. The dependency relationships include the dependent resources of each node, and the dependent resources include models;
[0184] The second acquisition module 930 is configured to, during the execution of the target workflow, acquire the dependent resources of the executed nodes in the directed acyclic graph as target dependent resources;
[0185] The second determination module 940 is configured to determine, based on the dependency relationships, whether the target dependent resources are depended on by the unexecuted nodes in the directed acyclic graph;
[0186] The processing module 950 is configured to release the target dependent resources when the target dependent resources are not depended on by the unexecuted nodes in the directed acyclic graph.
[0187] In an alternative embodiment, the first determination module 920 is further configured to:
[0188] Initialize a dependency mapping, which is used to illustrate each of the dependent resources and the consuming nodes corresponding to the dependent resources;
[0189] Traverse each node in the directed acyclic graph, and during the traversal, determine the dependent resources of the current node as the first dependent resources based on the input of the current node;
[0190] In the case where the first dependent resources are not in the dependency mapping, construct an output key of the first dependent resources and a set of consuming nodes corresponding to the output key in the dependency mapping, and add the current node to the set of consuming nodes. The output key includes a source node and an output index;
[0191] In the case where the first dependent resources are already in the dependency mapping, add the current node to the set of consuming nodes corresponding to the first dependent resources.
[0192] In an alternative embodiment, the apparatus 900 is further configured to:
[0193] When the execution of the first node is completed, traverse the set of consumer nodes corresponding to each of the output keys in the dependency mapping, where the first node is any node of the directed acyclic graph;
[0194] When the first node is included in the currently traversed set of consumer nodes, remove the first node from the currently traversed set of consumer nodes;
[0195] When the currently traversed set of consumer nodes is empty after the removal, determine that the dependency resources corresponding to the currently traversed set of consumer nodes are not depended on by the unexecuted nodes of the directed acyclic graph.
[0196] In an alternative embodiment, the processing module 950 is further configured to:
[0197] When the target dependency resources are not depended on by the unexecuted nodes of the directed acyclic graph, determine whether the video memory pressure of the current GPU video memory is less than a first pressure threshold;
[0198] When the video memory pressure is less than the first pressure threshold, skip the step of releasing the target dependency resources.
[0199] In an alternative embodiment, the apparatus 900 is further configured to:
[0200] When the video memory pressure is greater than or equal to the first pressure threshold, determine whether the video memory pressure is less than a second pressure threshold;
[0201] When the video memory pressure is less than the second pressure threshold, transfer the target dependency resources from the GPU video memory to the CPU memory;
[0202] When the video memory pressure is greater than or equal to the second pressure threshold, transfer the target dependency resources from the GPU video memory to the disk.
[0203] In an alternative embodiment, the apparatus 900 is further configured to:
[0204] Before executing the second node of the directed acyclic graph, determine the dependency resources corresponding to the second node as second dependency resources, where the second node is any node of the directed acyclic graph;
[0205] Check in sequence whether the second dependency resources are in the GPU video memory, the CPU memory, or the disk;
[0206] When the second dependent resource is in the CPU memory or the disk, load the second dependent resource into the GPU video memory.
[0207] In an alternative embodiment, the apparatus 900 is further configured to:
[0208] Intercept a target function from a third-party node, where the target function includes model loading, model unloading, and migration operation functions;
[0209] When the target function is intercepted, add the workflow corresponding to the target function to the directed acyclic graph of the target workflow.
[0210] Embodiment III
[0211] Figure 12 Schematically shows a hardware architecture diagram of a computer device 10000 suitable for implementing a computer resource processing method according to Embodiment III of the present application. In some embodiments, the computer device 10000 may be a terminal device such as a smart phone, a wearable device, a tablet computer, a personal computer, a vehicle-mounted terminal, a game console, a virtual device, a workbench, a digital assistant, a set-top box, a robot, etc. In other embodiments, the computer device 10000 may be a rack server, a blade server, a tower server, or a cabinet server (including an independent server or a server cluster composed of multiple servers). As Figure 12 shown, the computer device 10000 includes, but is not limited to: a memory 10010, a processor 10020, and a network interface 10030 that can communicate with each other through a system bus. Among them:
[0212] The memory 10010 includes at least one type of computer-readable storage medium. The readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the memory 10010 may be an internal storage module of the computer device 10000, such as the hard disk or memory of the computer device 10000. In other embodiments, the memory 10010 may also be an external storage device of the computer device 10000, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the computer device 10000. Of course, the memory 10010 may also include both the internal storage module and the external storage device of the computer device 10000. In this embodiment, the memory 10010 is generally used to store the operating system installed on the computer device 10000 and various application software, such as the program code of the computer resource processing method. In addition, the memory 10010 can also be used to temporarily store various types of data that have been output or will be output.
[0213] In some embodiments, the processor 10020 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other chips. The processor 10020 is generally used to control the overall operation of the computer device 10000, such as performing control and processing related to data interaction or communication with the computer device 10000. In this embodiment, the processor 10020 is used to run the program code stored in the memory 10010 or process data.
[0214] The network interface 10030 may include a wireless network interface or a wired network interface, which is generally used to establish a communication link between the computer device 10000 and other computer devices. For example, the network interface 10030 is used to connect the computer device 10000 to an external terminal through a network, and establish a data transmission channel and a communication link between the computer device 10000 and the external terminal. The network may be a wireless or wired network such as an enterprise intranet (Intranet), the Internet, Global System of Mobile communication (GSM for short), Wideband Code Division Multiple Access (WCDMA for short), 4G network, 5G network, Bluetooth, Wi-Fi, etc.
[0215] It should be noted that Figure 12 Only the computer device with components 10010 - 10030 is shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components may be implemented alternatively.
[0216] In this embodiment, the computer resource processing method stored in the memory 10010 may also be divided into one or more program modules and executed by one or more processors (such as the processor 10020) to complete the embodiments of the present application.
[0217] Embodiment 4
[0218] The embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the computer resource processing method in the embodiments are implemented.
[0219] In this embodiment, the computer-readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the computer-readable storage medium may be an internal storage unit of a computer device, such as the hard disk or memory of the computer device. In other embodiments, the computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the computer device. Of course, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the computer device. In this embodiment, the computer-readable storage medium is generally used to store the operating system installed on the computer device and various application software, such as the program code of the computer resource processing method in the embodiment. In addition, the computer-readable storage medium can also be used to temporarily store various data that have been output or will be output.
[0220] Embodiment Five
[0221] The embodiment of the present application also provides a computer program product, including a computer program, which when executed by a processor implements the method in the above embodiment.
[0222] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the embodiments of the present application can be implemented by a general-purpose computer device. They can be concentrated on a single computer device or distributed on a network composed of multiple computer devices. Optionally, they can be implemented by program codes executable by the computer device, so that they can be stored in a storage device and executed by the computer device. And in some cases, the steps shown or described can be executed in a different order from here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module to implement. Thus, the embodiments of the present application are not limited to any specific combination of hardware and software.
[0223] It should be noted that the above are only the preferred embodiments of the present application, and do not limit the patent protection scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, are equally included in the patent protection scope of the present application.
Claims
1. A method for processing computer resources, characterized in that, The method includes: Obtaining a directed acyclic graph of a target workflow; Determining the dependency relationships of each node in the directed acyclic graph, where the dependency relationships include the dependent resources of each node, and the dependent resources include models; During the execution of the target workflow, obtaining the dependent resources of the executed nodes in the directed acyclic graph as target dependent resources; Based on the dependency relationships, determining whether the target dependent resources are depended on by the unexecuted nodes in the directed acyclic graph; When the target dependent resources are not depended on by the unexecuted nodes in the directed acyclic graph, releasing the target dependent resources; The determining the dependency relationships of each node in the directed acyclic graph includes: Initializing a dependency mapping, which is used to illustrate each of the dependent resources and the consuming nodes corresponding to the dependent resources; Traversing each node in the directed acyclic graph, and during the traversal, determining the dependent resources of the current node as the first dependent resources based on the input of the current node; When the first dependent resources are not in the dependency mapping, constructing an output key of the first dependent resources and a set of consuming nodes corresponding to the output key in the dependency mapping, and adding the current node to the set of consuming nodes, where the output key includes a source node and an output index; When the first dependent resources are already in the dependency mapping, adding the current node to the set of consuming nodes corresponding to the first dependent resources.
2. The method according to claim 1, wherein The during the execution of the target workflow, obtaining the dependent resources of the executed nodes in the directed acyclic graph as target dependent resources, and based on the dependency relationships, determining whether the target dependent resources are depended on by the unexecuted nodes in the directed acyclic graph includes: When a first node is executed, traversing the set of consuming nodes corresponding to each output key in the dependency mapping, where the first node is any node in the directed acyclic graph; When the first node is included in the currently traversed set of consuming nodes, deleting the first node from the currently traversed set of consuming nodes; When the currently traversed set of consuming nodes is empty after deletion, determining that the dependent resources corresponding to the currently traversed set of consuming nodes are not depended on by the unexecuted nodes in the directed acyclic graph.
3. The method according to claim 1, wherein The when the target dependent resources are not depended on by the unexecuted nodes in the directed acyclic graph, releasing the target dependent resources includes: When the target dependent resources are not depended on by the unexecuted nodes in the directed acyclic graph, determining whether the current GPU video memory pressure is less than a first pressure threshold; When the video memory pressure is less than the first pressure threshold, skipping the step of releasing the target dependent resources.
4. The method according to claim 3, wherein The method further includes: When the video memory pressure is greater than or equal to the first pressure threshold, determining whether the video memory pressure is less than a second pressure threshold; When the video memory pressure is less than the second pressure threshold, transferring the target dependent resources from the GPU video memory to the CPU memory; When the video memory pressure is greater than or equal to the second pressure threshold, transfer the target dependent resource from the GPU video memory to the disk.
5. The method according to any one of claims 1 to 4, characterized in that The method further includes: Before executing the second node of the directed acyclic graph, determine the dependent resource corresponding to the second node as the second dependent resource, where the second node is any node of the directed acyclic graph; Check in sequence whether the second dependent resource is in the GPU video memory, CPU memory or disk; When the second dependent resource is in the CPU memory or the disk, load the second dependent resource into the GPU video memory.
6. The method according to any one of claims 1-4, characterized in that The method further includes: Intercept the target function from a third-party node, where the target function includes a model loading, model unloading and migration operation function; When the target function is intercepted, add the workflow corresponding to the target function to the directed acyclic graph of the target workflow.
7. A computer resource processing device, characterized in that, The device includes: A first acquisition module, configured to acquire a directed acyclic graph of a target workflow; A first determination module, configured to determine the dependency relationship of each node in the directed acyclic graph, where the dependency relationship includes the dependent resources of each node, and the dependent resources include models; A second acquisition module, configured to, during the execution of the target workflow, acquire the dependent resources of the executed nodes in the directed acyclic graph as target dependent resources; A second determination module, configured to determine based on the dependency relationship whether the target dependent resource is depended on by the unexecuted nodes of the directed acyclic graph; A processing module, configured to release the target dependent resource when the target dependent resource is not depended on by the unexecuted nodes of the directed acyclic graph; The first determination module is further configured to: Initialize a dependency mapping, where the dependency mapping is used to illustrate each of the dependent resources and the consumer nodes corresponding to the dependent resources; Traverse each node in the directed acyclic graph, and during the traversal, determine the dependent resource of the current node as the first dependent resource based on the input of the current node; When the first dependent resource is not in the dependency mapping, construct an output key of the first dependent resource and a set of consumer nodes corresponding to the output key in the dependency mapping, and add the current node to the set of consumer nodes, where the output key includes a source node and an output index; When the first dependent resource is already in the dependency mapping, add the current node to the set of consumer nodes corresponding to the first dependent resource.
8. A computer device, characterized in that, Includes: At least one processor; And A memory communicatively connected to the at least one processor; wherein: The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, Computer instructions are stored in the computer-readable storage medium, and when the computer instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Dependent task unloading system and method based on mobile edge computing
CN114980216A