Computer resource processing method and device

By analyzing the dependencies in the directed acyclic graph of the workflow, releasing unreliable model resources, solving the problems of memory overflow and long execution time in multi-model collaborative scenarios, and achieving more efficient resource management and system performance.

CN120104347AActive Publication Date: 2025-06-06SHANGHAI HODE INFORMATION TECH CO LTD

Patent Information

Application Number
CN202510585215.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-06-06
Estimated Expiration
2045-05-08

AI Technical Summary

Technical Problem

In the scenario where multiple artificial intelligence models work together, it is difficult for the prior art to effectively manage the loading and unloading of models, resulting in memory overflow or workflow execution time being too long.

Method used

By obtaining the directed acyclic graph of the target workflow, the dependencies of each node are determined, the dependencies of the executed nodes are obtained, and based on the dependencies, whether these resources are dependent on the unexecuted nodes. Unable to depend on these resources to avoid memory overflow and optimize workflow execution time.

Benefits of technology

It effectively avoids the problem of video memory overflow, while reasonably reducing the workflow execution time and improving system efficiency and response speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104347A_ABST
    Figure CN120104347A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a computer resource processing method, and relates to the field of electric digital data processing, and the computer resource processing method comprises the steps: obtaining a directed acyclic graph of a target workflow; determining a dependency relationship of each node in the directed acyclic graph, the dependency relationship including a dependency resource of each node, and the dependency resource including a model; in the execution process of the target workflow, obtaining a dependency resource of an executed node in the directed acyclic graph as a target dependency resource; determining whether the target dependency resource is depended by an unexecuted node of the directed acyclic graph or not based on the dependency relationship; and releasing the target dependency resource under the condition that the target dependency resource is not depended by the unexecuted nodes of the directed acyclic graph. According to the technical scheme, the problem of video memory overflow in the multi-model cooperation process can be effectively avoided, meanwhile, the workflow execution time can be reasonably shortened, and the system efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of computer technology, and in particular, to a computer resource processing method, apparatus, computer equipment, computer-readable storage medium, and computer program product. Background Art

[0002] In the field of artificial intelligence applications, it is usually involved in scenarios where multiple artificial intelligence models (such as deep learning models) work together. In this scenario, loading and unloading artificial intelligence models is a key operation that directly affects the utilization efficiency and execution performance of system resources. There are currently two main model management strategies: 1. Continuous retention strategy: once the model is loaded into the GPU memory, it is not actively released so that it can be reused in subsequent steps; 2. Timely release strategy: the model is released from the GPU memory immediately after the inference is completed, and reloaded each time it is used.

[0003] However, although the first strategy can reduce the time overhead of repeated model loading, it is easy to cause memory accumulation and cause memory overflow (OOM) errors; although the second strategy can effectively avoid the memory overflow problem, frequent model loading and unloading operations will significantly increase the workflow execution time and reduce the overall efficiency of the system.

[0004] It should be noted that the above content is not necessarily prior art, nor is it intended to limit the scope of patent protection of this application. Summary of the invention

[0005] The embodiments of the present application provide a computer resource processing method, apparatus, computer device, computer-readable storage medium, and computer program product to solve or alleviate one or more of the technical problems raised above.

[0006] One aspect of an embodiment of the present application provides a computer resource processing method, the method comprising: Get the directed acyclic graph of the target workflow; Determine the dependency relationship of each node in the directed acyclic graph, wherein the dependency relationship includes the dependent resources of each node, and the dependent resources include a model; During the execution of the target workflow, obtaining dependent resources of executed nodes in the directed acyclic graph as target dependent resources; Determining whether the target dependent resource is dependent on an unexecuted node of the directed acyclic graph based on the dependency relationship; When the target dependent resource is not dependent on any unexecuted node of the directed acyclic graph, the target dependent resource is released.

[0007] Optionally, determining the dependency relationship of each node in the directed acyclic graph includes: Initialize a dependency map, where the dependency map is used to describe each of the dependent resources and the consumption nodes corresponding to the dependent resources; Traversing each node in the directed acyclic graph, and in the traversal process, determining a dependent resource of the current node as a first dependent resource based on an input of the current node; If the first dependent resource is not in the dependency map, constructing an output key of the first dependent resource and a consumer node set corresponding to the output key in the dependency map, and adding the current node to the consumer node set, wherein the output key includes a source node and an output index; In the case where the first dependent resource is already in the dependency map, the current node is added to a set of consumer nodes corresponding to the first dependent resource.

[0008] Optionally, during the execution of the target workflow, obtaining dependent resources of executed nodes in the directed acyclic graph as target dependent resources, and determining whether the target dependent resources are dependent on unexecuted nodes of the directed acyclic graph based on the dependency relationship, includes: When the first node is executed, traverse the set of consumer nodes corresponding to each output key in the dependency map, where the first node is any node of the directed acyclic graph; In the case where the first node is included in the consumer node set currently traversed, deleting the first node from the consumer node set currently traversed; When the currently traversed consumer node set is empty after the deletion, it is determined that the dependent resources corresponding to the currently traversed consumer node set are not dependent on the unexecuted nodes of the directed acyclic graph.

[0009] Optionally, when the target dependent resource is not dependent on an unexecuted node of the directed acyclic graph, releasing the target dependent resource includes: In a case where the target dependent resource is not dependent on an unexecuted node of the directed acyclic graph, determining whether a video memory pressure of the current GPU video memory is less than a first pressure threshold; When the video memory pressure is less than the first pressure threshold, the step of releasing the target dependent resource is skipped.

[0010] Optionally, the method further comprises: When the video memory pressure is greater than or equal to the first pressure threshold, determining whether the video memory pressure is less than a second pressure threshold; When the video memory pressure is less than the second pressure threshold, transferring the target dependent resource from the GPU video memory to the CPU memory; When the video memory pressure is greater than or equal to the second pressure threshold, the target dependent resource is transferred from the GPU video memory to a disk.

[0011] Optionally, the method further comprises: Before executing a second node of the directed acyclic graph, determining a dependent resource corresponding to the second node as a second dependent resource, wherein the second node is any node of the directed acyclic graph; Checking in sequence whether the second dependent resource is in the GPU video memory, the CPU memory or the disk; In the case that the second dependent resource is in the CPU memory or the disk, the second dependent resource is loaded into the GPU video memory.

[0012] Optionally, the method further comprises: Intercepting target functions from third-party nodes, wherein the target functions include model loading, model unloading and migration operation functions; When the target function is intercepted, the workflow corresponding to the target function is added to the directed acyclic graph of the target workflow.

[0013] Another aspect of an embodiment of the present application provides a computer resource processing device, the device comprising: A first acquisition module is used to acquire a directed acyclic graph of a target workflow; A first determination module is used to determine the dependency relationship of each node in the directed acyclic graph, wherein the dependency relationship includes the dependent resources of each node, and the dependent resources include a model; A second acquisition module is used to acquire, during the execution of the target workflow, dependent resources of executed nodes in the directed acyclic graph as target dependent resources; A second determination module, configured to determine, based on the dependency relationship, whether the target dependent resource is dependent on an unexecuted node of the directed acyclic graph; The processing module is used to release the target dependent resource when the target dependent resource is not dependent on the unexecuted node of the directed acyclic graph.

[0014] Another aspect of an embodiment of the present application provides a computer device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein: the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described above.

[0015] Another aspect of an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the method described above is implemented.

[0016] Another aspect of an embodiment of the present application provides a computer program product, including a computer program, which implements the method described above when executed by a processor.

[0017] The above technical solution adopted in the embodiment of the present application may have the following advantages: By obtaining the directed acyclic graph of the target workflow, the dependency relationship of each node in the directed acyclic graph is determined, and the dependency relationship includes the dependent resources of each node, and the dependent resources include models; during the execution of the target workflow, the dependent resources of the executed nodes in the directed acyclic graph are obtained as the target dependent resources, and based on the dependency relationship, it is determined whether the target dependent resources are dependent on the unexecuted nodes of the directed acyclic graph. When the target dependent resources are not dependent on the unexecuted nodes of the directed acyclic graph, the target dependent resources are released. By analyzing the dependency relationship of each node in the directed acyclic graph of the workflow, it is determined according to the dependency relationship whether the resources (including models) in the system are dependent on the unexecuted nodes, and the resources are released when they are not dependent. This can effectively avoid the problem of video memory overflow in scenarios such as multi-model collaboration, and at the same time can reasonably reduce the workflow execution time, improve the response speed of requests and system efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The accompanying drawings exemplarily illustrate the embodiments and constitute a part of the specification, and together with the text description of the specification, are used to explain the exemplary implementation of the embodiments. The embodiments shown are for illustrative purposes only and do not limit the scope of the claims. In all drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0019] Figure 1 The operating environment diagram of the computer resource processing method according to the first embodiment of the present application is schematically shown; Figure 2 The flowchart of the computer resource processing method according to the first embodiment of the present application is schematically shown; Figure 3 Schematically shows Figure 2 Flow chart of sub-steps of step S202; Figure 4 Schematically shows Figure 2 Flow chart of sub-steps of step S204 and step S206; Figure 5 Schematically shows Figure 2 Flow chart of sub-steps of step S208; Figure 6 The newly added process of the computer resource processing method according to the first embodiment of the present application is schematically shown; Figure 7 Another newly added process of the computer resource processing method according to the first embodiment of the present application is schematically shown; Figure 8 Schematically illustrates another newly added process of the computer resource processing method according to the first embodiment of the present application; Fig. 9 The system architecture diagram of the computer resource processing method according to the first embodiment of the present application is schematically shown; Fig.10 The following is a schematic diagram showing an example of a timing flow of a computer resource processing method according to the first embodiment of the present application; Fig.11 A block diagram schematically shows a computer resource processing device according to the second embodiment of the present application; and Fig.12 The hardware architecture diagram of the computer device according to the third embodiment of the present application is schematically shown. DETAILED DESCRIPTION

[0020] In order to make the purpose, technical solutions and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application.

[0021] It should be noted that the descriptions involving "first", "second", etc. in the embodiments of the present application are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In addition, the technical solutions between the various embodiments can be combined with each other, but they must be based on the ability of ordinary technicians in the field to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such combination of technical solutions does not exist and is not within the scope of protection required by this application.

[0022] In the description of the present application, it should be understood that the numerical labels before the steps do not indicate the order in which the steps are executed, but are only used to facilitate the description of the present application and to distinguish each step, and therefore should not be understood as a limitation on the present application.

[0023] First, the following terms are explained: Workflow Engine: A software component or system used to define, manage and automatically execute workflows. It achieves automated management of business processes by breaking down business processes into a series of executable task nodes and driving task flow according to preset rules and logic (such as sequence, conditional branches, parallel processing, etc.).

[0024] Directed Acyclic Graph (DAG): The basic data structure of a workflow that represents the dependencies between nodes and ensures that nodes are executed in the correct order without forming circular dependencies.

[0025] Video Memory Overflow: It is a common technical problem in computer graphics, deep learning or high-performance computing. It refers to the abnormal state in which the video RAM (VRAM) of the graphics processing unit (GPU) cannot accommodate the data or computing tasks that need to be processed, resulting in the exhaustion of memory resources.

[0026] Community Node: A workflow node created by third-party developers.

[0027] Function Interception: It is a programming technique that refers to inserting custom logic before or after the target function is called (or during execution) to monitor, modify or control the behavior of the function without directly modifying the source code of the target function. This technique is often used in debugging, logging, performance analysis, permission control, input and output filtering and other scenarios. The core goal is to expand or change the program function without invading the original code.

[0028] Dependency Graph: A directed graph used to describe the dependencies between components, modules, tasks, data or variables in a system. It is often used in software engineering, project management, compiler theory, data processing and other fields. The core idea is to visualize the dependencies between elements through a graph structure, helping to analyze system complexity, execution order, conflicts or circular dependencies and other issues.

[0029] Hook Function: A callback function triggered by a specific event, used to implement custom logic such as resource management.

[0030] Secondly, in order to facilitate those skilled in the art to understand the technical solutions provided in the embodiments of the present application, the relevant technologies are described below: In scenarios where multiple AI models work together, loading and unloading AI models is a critical operation that directly affects the utilization efficiency and execution performance of system resources. Currently, there are two main model management strategies: 1. Continuous retention strategy: once the model is loaded into the GPU memory, it will not be released actively; 2. Timely release strategy: each time the model completes reasoning, it will be released from the GPU memory immediately and reloaded the next time it is used.

[0031] However, although the first strategy can reduce the time overhead of repeated model loading, it is easy to cause memory accumulation and memory overflow errors; although the second strategy can avoid memory overflow, it will significantly increase the workflow execution time and reduce system efficiency.

[0032] To this end, the embodiment of the present application provides a computer resource processing technical solution. In this technical solution, by analyzing the dependency relationship of each node in the directed acyclic graph of the workflow, it is determined whether the resources (including models) in the system are dependent on the unexecuted nodes according to the dependency relationship, and the resources are released when they are not dependent. This can effectively avoid the problem of video memory overflow in scenarios such as multi-model collaboration, and at the same time can reasonably reduce the workflow execution time and improve system efficiency. See below for details.

[0033] Finally, for ease of understanding, an exemplary operating environment is provided below.

[0034] like Figure 1 As shown, the environment diagram includes a service platform 2, a network 4, and a client 6, wherein: The service platform 2 may be composed of a single or multiple computing devices. The multiple computing devices may include virtualized computing instances. Virtualized computing instances may include virtual machines, such as simulations of computer systems, operating systems, servers, etc. The computing device may load a virtual machine based on a virtual image and / or other data defining specific software (e.g., operating system, dedicated application, server) for simulation. As the demand for different types of processing services changes, different virtual machines may be loaded and / or terminated on one or more computing devices. A hypervisor may be implemented to manage the use of different virtual machines on the same computing device.

[0035] The service platform 2 may be configured to communicate with the client 6 or the like via a network 4. The network 4 includes various network devices, such as routers, switches, multiplexers, hubs, modems, bridges, repeaters, firewalls, proxy devices, and / or the like. The network 4 may include physical links, such as coaxial cable links, twisted pair cable links, optical fiber links, combinations thereof, and the like, or wireless links, such as cellular links, satellite links, Wi-Fi links, and the like.

[0036] The service platform 2 can provide storage, reading, writing, querying, deleting and other services, such as providing drawing, video making and other services for the client.

[0037] The client 6 may be an electronic device running an operating system such as Windows, Android™ or iOS, such as a smart phone, a tablet device, a laptop computer, a virtual reality device, a gaming device, a set-top box, a vehicle terminal, or a smart TV. Based on the above operating system, various applications may be run, such as drawing, video making, and the like.

[0038] The client 6 may provide / configure a user access page for manipulating the service platform 2 or uploading an object, etc.

[0039] It should be noted that the above devices are exemplary, and the number and type of devices are adjustable in different scenarios or according to different needs.

[0040] The technical solutions of the present application are described below through multiple embodiments. It should be noted that these embodiments can be implemented in a variety of different forms and should not be construed as being limited to the embodiments described here.

[0041] Embodiment 1 Figure 2 The flowchart of the computer resource processing method according to the first embodiment of the present application is schematically shown. It should be noted that the execution subject of the computer resource processing method of the embodiment of the present application can be Figure 1 A service platform, or more specifically, a computing node in the service platform.

[0042] like Figure 2 As shown, the computer resource processing method may include steps S200 to S208, wherein: Step S200: Obtain a directed acyclic graph of the target workflow.

[0043] Step S202, determining the dependency relationship of each node in the directed acyclic graph, the dependency relationship includes the dependent resources of each node, and the dependent resources include the model.

[0044] Step S204: during the execution of the target workflow, dependent resources of executed nodes in the directed acyclic graph are obtained as target dependent resources.

[0045] Step S206: determining whether the target dependent resource is dependent on an unexecuted node of the directed acyclic graph based on the dependency relationship.

[0046] Step S208 : when the target dependent resource is not dependent on any unexecuted node of the directed acyclic graph, release the target dependent resource.

[0047] The computer resource processing method provided in this embodiment obtains the directed acyclic graph of the target workflow, determines the dependency relationship of each node in the directed acyclic graph, and the dependency relationship includes the dependent resources of each node, and the dependent resources include models; during the execution of the target workflow, obtains the dependent resources of the executed nodes in the directed acyclic graph as the target dependent resources, and determines whether the target dependent resources are dependent on the unexecuted nodes of the directed acyclic graph based on the dependency relationship. When the target dependent resources are not dependent on the unexecuted nodes of the directed acyclic graph, the target dependent resources are released. The dependency relationship of each node in the directed acyclic graph of the workflow can be analyzed, and the resources (including models) in the system can be determined according to the dependency relationship. When they are not dependent, the resources are released, which can effectively avoid the problem of video memory overflow in scenarios such as multi-model collaboration, and at the same time can reasonably reduce the workflow execution time, improve the request response speed and system efficiency.

[0048] The following combination Figure 2 , each step in steps S200~S208 and other optional steps are explained in detail.

[0049] Step S200 , obtain the directed acyclic graph of the target workflow.

[0050] The target workflow can be a specific or arbitrary workflow in the system. The directed acyclic graph of the target workflow can be constructed or parsed according to the specific business logic (such as code, configuration file, and visual metadata) of the target workflow. In practical applications, the directed acyclic graph of the target workflow can be completed by a workflow engine. When obtaining the directed acyclic graph of the target workflow, it can be obtained through the workflow engine.

[0051] Step S202 , determine the dependency relationship of each node in the directed acyclic graph, the dependency relationship includes the dependent resources of each node, and the dependent resources include the model.

[0052] Specifically, the dependent resources of each node can be determined based on the input of each node in the directed acyclic graph and the connection relationship of each node, wherein the dependent resources can be the output of a certain node. For example, if a certain node a calls the A model to stylize the image of the previous node, then the dependent resources generated by the node are the A model and the image of the previous node, and its output is a stylized image; if the input of a certain node b is the stylized image, then the dependent resource of node b can be the stylized image; for another example, if node c loads the A model, then the A model is the output of node c. By performing the same analysis on all nodes in the directed acyclic graph, the dependency relationship of each node can be determined. Among them, the dependency relationship can be specifically represented in the form of a dependency list, a dependency tree, a dependency graph, and a dependency network, which are not limited here.

[0053] Step S204 ,During the execution of the target workflow, the dependent resources of the executed nodes in the directed acyclic graph are obtained as the target dependent resources.

[0054] Specifically, after a certain node is executed, its corresponding dependent resources are obtained as target dependent resources to determine whether the target dependent resources are depended on by other unexecuted nodes. For example, if node d has been executed and the dependent resource of node d is model B, model B can be used as the target dependent resource. Optionally, after a certain node is executed, the dependent resources of all executed nodes can be obtained as target dependent resources. For example, if nodes a, b, c, and d of a directed acyclic graph have been executed, the dependent resources corresponding to nodes a, b, c, and d can be obtained as target dependent resources, that is, a comprehensive check is performed on the dependent resources of all executed nodes.

[0055] Step S206 , based on the dependency relationship, determine whether the target dependent resource is dependent on the unexecuted node of the directed acyclic graph.

[0056] Specifically, it is possible to determine whether the target dependent resource is dependent on the unexecuted node based on the dependency relationship obtained by previous analysis. For example, if it is found through the dependency relationship that the target dependent resource B model does not need any unexecuted nodes, then it is determined that the B model is not dependent on the unexecuted nodes; if it is found through the dependency relationship that the B model is dependent on the unexecuted node e, then it is determined that the B model is dependent on the unexecuted node.

[0057] Step S208 , when the target dependent resource is not depended on by the unexecuted nodes of the directed acyclic graph, the target dependent resource is released.

[0058] Specifically, if it is determined that the target dependent resource is not dependent on the unexecuted node, the target dependent resource can be released to free up space, reduce GPU memory accumulation, and avoid the problem of memory overflow. For example, if the target dependent resource B model is not dependent on the unexecuted node, the B model can be released. Of course, when it is determined that the target dependent resource is not dependent on the unexecuted node, other conditions can be judged (such as whether the video memory is sufficient). Only when other conditions are met, the target dependent resource is released, otherwise the target dependent resource is retained. For example, a conditional judgment can be made on whether the video memory is sufficient. If the video memory is insufficient, the target dependent resource is released. If the video memory is sufficient, the target dependent resource is retained for possible subsequent use.

[0059] In an optional embodiment, if Figure 3 As shown, in step S202, determining the dependency relationship of each node in the directed acyclic graph may include: Step S300, initializing dependency mapping, where the dependency mapping is used to describe each dependent resource and the consumption node corresponding to the dependent resource.

[0060] Step S302, traversing each node in the directed acyclic graph, and in the traversal process, determining a dependent resource of the current node as a first dependent resource based on an input of the current node.

[0061] Step S304, when the first dependent resource is not in the dependency map, construct the output key of the first dependent resource and the consumer node set corresponding to the output key in the dependency map, and add the current node to the consumer node set, where the output key includes the source node and the output index.

[0062] Step S306: When the first dependent resource is already in the dependency map, add the current node to the set of consumer nodes corresponding to the first dependent resource.

[0063] Specifically, a blank dependency map can be initialized, and then the traversal starts from the first node of the directed acyclic graph. For each traversed node, its input is analyzed to determine the dependency resource corresponding to the current traversed node as the first dependency resource. If the first dependency resource is not in the dependency map, the output key of the first dependency resource and the consumption node set corresponding to the output key are constructed in the dependency map, and the current node is added to the currently constructed consumption node set; if the first dependency resource is already in the dependency map, the current node is added to the consumption node set corresponding to the output key of the first dependency resource. When all nodes are traversed, a complete dependency map can be obtained. For example, if the first dependency resource corresponding to the current node is model C, model C is the first output of node a, and the first dependency resource is not in the dependency map, the output key of model C can be constructed in the dependency map as (node ​​a, 1), the consumption node set is empty, and then the current node is added to the consumption node set of model C.

[0064] In this embodiment, by initializing the dependency mapping, each node of the directed acyclic graph is traversed. During the traversal process, the dependent resource is determined as the first dependent resource based on the input of the current node. If the first dependent resource is not in the dependency mapping, the output key of the first dependent resource and the corresponding consumer node set are constructed in the dependency mapping, and the current node is added to the consumer node set; if the first dependent resource is already in the dependency mapping, the current node is added to the corresponding consumer node. This can effectively construct the dependency relationship of each node in the directed acyclic graph, which is convenient for the subsequent judgment of whether to release the resources.

[0065] In an optional embodiment, in step S204 and step S206, that is, during the execution of the target workflow, the dependent resources of the executed nodes in the directed acyclic graph are obtained as the target dependent resources, and it is determined based on the dependency relationship whether the target dependent resources are dependent on the unexecuted nodes of the directed acyclic graph, such as Figure 4 As shown, it may include: Step S400, when the first node is executed, traverse the set of consumer nodes corresponding to each output key in the dependency map, where the first node is any node of the directed acyclic graph.

[0066] Step S402: when the first node is included in the currently traversed consumer node set, the first node is deleted from the currently traversed consumer node set.

[0067] Step S404, when the currently traversed consumer node set is empty after the deletion, it is determined that the dependent resources corresponding to the currently traversed consumer node set are not dependent on the unexecuted nodes of the directed acyclic graph.

[0068] Specifically, when the first node (any node) is executed, the consumption node set corresponding to each output key in the dependency map is traversed, and when the first node is included in the currently traversed consumption node set, the first node is deleted from the currently traversed consumption node set; after deletion, the currently traversed consumption node set is judged, and if the currently traversed consumption node set is empty after deletion, it is determined that the dependent resources corresponding to the currently traversed consumption node set are not dependent on the unexecuted nodes of the directed acyclic graph; if the currently traversed consumption node set is not empty after deletion, it is determined that the dependent resources corresponding to the currently traversed consumption node set are dependent on the unexecuted nodes of the directed acyclic graph. For example, after the first node is executed, the consumption node set in the dependency map is traversed, and if the consumption node set of model A is traversed to include the first node, the first node is deleted from the consumption node set of model A; if the consumption node set corresponding to model A is empty after deletion, it is determined that model A is not dependent on the unexecuted nodes. Among them, a hook function can be added after the node in the directed acyclic graph, so that after executing a node, the traversal and subsequent operations are started through the hook function. In practical applications, in order to facilitate the release of resources that are not dependent on unexecuted nodes, the dependent resources can be marked as releasable or deletable when it is determined that the dependent resources corresponding to the currently traversed consumer node set are not dependent on the unexecuted nodes. Optionally, the dependent resources in the dependency map can also be reference counted according to the dependent nodes. After a reference node is executed, the corresponding number of dependent resources is reduced by 1. Finally, it can be determined whether the dependent resource is dependent on the unexecuted node based on whether the reference count of the dependent resource is zero.

[0069] In this embodiment, when any node in the directed acyclic graph is executed, the consumer node set in the dependency map is traversed, the executed node is deleted from the consumer node set, and then whether the corresponding dependent resource is dependent on the unexecuted node is determined based on whether the deleted consumer node set is empty. This can effectively determine whether the resources in the system still need to run in the GPU video memory, so that the resources that can be released can be accurately determined.

[0070] In an optional embodiment, in step S208, when the target resource is not dependent on any unexecuted node in the directed acyclic graph, the target dependent resource is released, such as Figure 5 As shown, it may include: Step S500: When the target dependent resource is not dependent on an unexecuted node of the directed acyclic graph, determine whether the current GPU memory pressure is less than a first pressure threshold.

[0071] Step S502: when the video memory pressure is less than the first pressure threshold, skip the step of releasing the target dependent resource.

[0072] The first pressure threshold may be a value used to indicate whether the GPU video memory is sufficient, that is, if the video memory pressure is less than the first pressure threshold, it is said that the video memory is sufficient; if the video memory pressure is greater than or equal to the first pressure threshold, it is said that the video memory is insufficient. The setting of the first pressure threshold may be determined according to specific application scenarios, hardware configurations, and strategies for using GPU video memory, and no specific restrictions are made here. Among them, the video memory pressure may be comprehensively determined based on indicators such as the absolute size of the used video memory and the percentage of the used video memory in the total video memory.

[0073] Specifically, when it is determined that the target dependent resource is not dependent on the unexecuted nodes of the directed acyclic graph, it is further determined whether the current GPU video memory pressure is less than the first pressure threshold; when the video memory pressure is less than the first pressure threshold, it is determined that the current video memory is sufficient, then the step of releasing the target dependent resource can be skipped, and the target dependent resource can be temporarily retained in the GPU video memory.

[0074] In this embodiment, when the target dependent resource is not dependent on any unexecuted node of the directed acyclic graph, it is determined whether the current GPU memory pressure is less than a first pressure threshold. When the memory pressure is less than the first pressure threshold, the step of releasing the target dependent resource is skipped. When the memory is sufficient, the resource can be temporarily retained in the GPU memory so that it can be directly used when needed later without reloading, thereby reducing start-stop overhead.

[0075] In an optional embodiment, if Figure 6 As shown, the computer resource processing method of the embodiment of the present application may also include: Step S600: when the video memory pressure is greater than or equal to the first pressure threshold, determine whether the video memory pressure is less than the second pressure threshold.

[0076] Step S602: when the video memory pressure is less than the second pressure threshold, the target dependent resource is transferred from the GPU video memory to the CPU memory.

[0077] Step S604: when the video memory pressure is greater than or equal to the second pressure threshold, the target dependent resource is transferred from the GPU video memory to the disk.

[0078] The second pressure threshold can be a value used to indicate whether the video memory pressure is moderate, that is, if the video memory pressure is less than the second pressure threshold, it means that although the video memory is not sufficient, the video memory pressure is still moderate, which can be understood as medium video memory pressure; if the video memory pressure is greater than or equal to the second threshold, it means that the video memory pressure is high. Similarly, the setting of the second pressure threshold can be determined according to the specific application scenario, hardware configuration, and the strategy for the use of GPU video memory, and no specific restrictions are made here. Among them, the second pressure threshold is greater than the first pressure threshold.

[0079] Specifically, when the video memory pressure is greater than or equal to the first pressure threshold, it is further determined whether the video memory pressure is less than the second pressure threshold; if the video memory pressure is less than the second pressure threshold, it means that the video memory pressure is still at a medium level, and the target dependent resources can be transferred from the GPU video memory to the CPU memory, which can be specifically releasing the target dependent resources from the GPU video memory and migrating the running location of the target dependent resources to the CPU memory; if the video memory pressure is greater than or equal to the second pressure threshold, it means that the video memory pressure is high, and the target dependent resources can be transferred from the GPU video memory to the disk, which can be specifically releasing the target dependent resources from the GPU video memory and migrating the running location of the target dependent resources to the disk.

[0080] In this embodiment, when the video memory pressure is greater than or equal to the first pressure threshold, it is determined whether the video memory pressure is less than the second pressure threshold. When the video memory pressure is less than the second pressure threshold, the target dependent resources are transferred from the GPU video memory to the CPU memory; when the video memory pressure is greater than or equal to the second pressure threshold, the target dependent resources are transferred from the GPU video memory to the disk. When the video memory pressure is medium, the resources can be migrated to the CPU memory, and subsequent loading from the CPU memory can speed up the response speed; at the same time, when the pressure is high, the resources can also be migrated to the disk to reduce the video memory pressure, thereby realizing multi-level caching under different video memory pressures and improving the degree of refinement of resource management.

[0081] In an optional embodiment, if Figure 7 As shown, the computer resource processing method of the embodiment of the present application may also include: Step S700, before executing a second node of the directed acyclic graph, determining a dependent resource corresponding to the second node as a second dependent resource, wherein the second node is any node of the directed acyclic graph.

[0082] Step S702, checking in sequence whether the second dependent resource is in the GPU video memory, the CPU memory or the disk.

[0083] Step S704: when the second dependent resource is in the CPU memory or the disk, load the second dependent resource into the GPU memory.

[0084] Among them, by adding a hook function in front of all nodes of the directed acyclic graph, the content of step S700 and subsequent steps can be executed before the node is executed.

[0085] Specifically, before executing the second node of the directed acyclic graph, determine the dependent resource corresponding to the second node as the second dependent resource; then, check in turn whether the second dependent resource is in the GPU memory, the CPU memory or the disk; when the second dependent resource is in the GPU memory, the second dependent resource can be used directly; when the second dependent resource is in the CPU memory or the disk, load the second dependent resource into the GPU memory for use.

[0086] In this embodiment, before executing any node of the directed acyclic graph, the dependent resources corresponding to the current node are determined, and the corresponding dependent resources are checked in turn to see whether they are in the GPU video memory, CPU memory or disk. If the corresponding dependent resources are in the CPU memory or disk, the corresponding dependent resources are loaded into the GPU video memory. The resources can be loaded from the appropriate cache level, and the corresponding accuracy is prepared in advance for the execution of the node, thereby improving the efficiency of the node execution.

[0087] In an optional embodiment, if Figure 8 As shown, the computer resource processing method of the embodiment of the present application may also include: Step S800, intercepting target functions from third-party nodes, where the target functions include model loading, model unloading and migration operation functions.

[0088] Step S802: When the target function is intercepted, the workflow corresponding to the target function is added to the directed acyclic graph of the target workflow.

[0089] Specifically, function interception can be set in the access path of third-party nodes (such as community nodes). When the functions of third-party nodes include functions for model loading, model unloading, and model migration, they are determined as target functions for interception; then the workflows corresponding to these target functions are parsed, and the workflows corresponding to the target functions are added to the directed acyclic graph of the target workflow. Subsequently, these corresponding models can be included in the dependency relationship for management, thereby achieving the purpose of unified resource management of the models corresponding to the third-party nodes. Optionally, in order to accurately manage these model resources, the model identifier can be generated based on multiple factors such as model source, structural characteristics, and initialization parameters. When one factor is different, the generated model identifier is also different.

[0090] In this embodiment, by intercepting target functions such as model loading, model unloading and migration operations from third-party nodes, the intercepted target functions are added to the directed acyclic graph of the target workflow. The corresponding model management can be included in the unified resource management without modifying the target functions of the third-party nodes, thereby achieving resource optimization to a greater extent.

[0091] In order to make this application easier to understand, the following Fig. 9 and Fig.10 An exemplary application is provided. Fig. 9 is a system architecture diagram corresponding to the computer resource processing method of the embodiment of the present application, Fig.10 The figure is an example of a timing flow chart of a computer resource processing method.

[0092] 1. If Fig. 9 As shown, the system mainly includes a workflow engine executor, a model manager and an engine optimizer.

[0093] 1. Workflow engine executor: responsible for parsing and executing DAG workflows. This module receives the workflow graph defined by the user and executes the computing tasks of each node according to the dependency relationship between nodes.

[0094] Specifically, the workflow engine executor can achieve the following functions: ⑴ Sort the workflow DAG topology to determine the order of node execution; ⑵Data transmission and parameter mapping between nodes; ⑶Node execution status management and exception handling.

[0095] 2. Model Manager: Responsible for model loading, unloading and cache management, and implements a three-level cache architecture, namely: ⑴L0 cache: GPU video memory, used to directly execute model reasoning, with the fastest access speed; ⑵L1 cache: CPU memory, used to temporarily store models that are not in use but may be called soon; ⑶L2 cache: disk storage, which saves the original weight file of the model.

[0096] The model manager provides a unified model loading and unloading interface for adapted standard nodes, and intercepts model operations of community nodes that have not completed adaptation through a hook mechanism, bringing them under unified management.

[0097] The key mechanisms of the model manager are: ⑴ Model loading optimization: take the optimal loading path according to the location of the model in the three-level cache; ⑵Model offloading strategy: perform precise offloading based on dependency analysis results; ⑶ Cache level conversion: dynamically adjust the storage location of the model between the three-level caches according to the workflow execution status and video memory pressure; ⑷Model status tracking: real-time monitoring of the model’s loading status, usage status, and cache location.

[0098] In terms of implementation, the model manager will load the model from the appropriate cache level according to the workflow node requirements. First, the L0 cache is checked. If it does not exist, the L1 and L2 caches are checked in turn. This mechanism can effectively reduce the overhead of repeated model loading.

[0099] 3. Engine Optimizer: It is a bridge connecting the workflow engine executor and the model manager, responsible for realizing intelligent resource management of workflows. This module achieves resource optimization through three key mechanisms: ⑴ Non-standard node interception: intercept the model loading and unloading logic of the community node and redirect it to the model manager for unified management. The interception mechanism does not modify the original node code, but dynamically takes over the resource management operation at runtime through function hook technology.

[0100] ⑵Workflow dependency analysis: Before the workflow is executed, the dependency graph between nodes is analyzed and constructed, and each node output is recorded by which subsequent nodes reference it. This information is the basis for dynamic pruning (resource release).

[0101] ⑶ Node execution hook: implement hook functions before and after each node execution in the workflow for resource preloading and automatic pruning. The post-node execution hook checks the dependency graph, identifies resources that are no longer needed, and triggers the uninstallation operation.

[0102] Dependency analysis and dynamic pruning in the embodiment of the present application are the core of resource optimization, which mainly includes the following steps: ⑴Static dependency analysis: Before the workflow is executed, perform a depth-first traversal of the DAG to identify the input-output dependency between nodes. For each node output, record all subsequent nodes that depend on the output and build a complete dependency relationship (dependency graph).

[0103] ⑵ Dynamic execution tracking: During the workflow execution process, the node execution status is continuously updated and the completed nodes are marked. Combined with the dependency graph information, the active status of each resource is calculated in real time.

[0104] ⑶ Intelligent pruning decision: After the node is executed, check the dependencies of its input and output. If an input or output is no longer referenced by an unexecuted node, mark the resources related to the input or output as releasable. Decide whether to perform the actual release operation based on resource characteristics and system status.

[0105] (4) Adaptive pruning strategy: When system resources are sufficient, pruning will be skipped to increase the probability that the model cache will be used by the next workflow; when resources are tight, pruning will be performed more aggressively.

[0106] 2. If Fig.10 As shown, the computer resource processing method may generally include the following processes: 1. The user submits the workflow DAG; 2. The workflow engine requests dependency analysis from the engine prioritizer, and the engine optimizer performs dependency analysis and builds a dependency graph; 3. The engine prioritizer returns the constructed dependency graph to the workflow engine; 4. For each node to be executed in the DAG, the workflow engine sends a request to the engine optimizer through the pre-execution hook, and the engine optimizer notifies the model manager to prepare model resources; 5. The model manager first determines whether the required model is in the L0 cache. If so, it is used directly; if in the L1 cache, it is loaded to the GPU; if in the L2 cache, the model is read and loaded to the GPU; 6. Resources are ready and preliminary preparations are completed.

[0107] 7. After the node is executed, the engine optimizer is used to update the dependency status through the post-execution hook, make pruning decisions, and request the model manager to release the specified resources; 8. The model manager determines whether the current video memory is sufficient. If sufficient, it skips the release of resources and keeps the resources in the L0 cache. If the video memory pressure is medium, it transfers the resources to the L1 cache. If the video memory pressure is high, it transfers the resources to the L2 cache. 9. Resource release completed; 10. Post-processing is completed; 11. After the workflow is executed, the workflow engine returns the workflow results to the user.

[0108] Embodiment 2 Fig.11 The block diagram of the computer resource processing device according to the second embodiment of the present application is schematically shown. The device can be divided into one or more program modules, one or more program modules are stored in a storage medium, and are executed by one or more processors to complete the embodiment of the present application. The program module referred to in the embodiment of the present application refers to a series of computer program instruction segments that can complete specific functions. The following description will specifically introduce the functions of each program module in this embodiment. Fig.11 As shown, the device 900 may include: a first acquisition module 910, a first determination module 920, a second acquisition module 930, a second determination module 940 and a processing module 950, wherein: A first acquisition module 910 is used to acquire a directed acyclic graph of a target workflow; A first determination module 920, configured to determine the dependency relationship of each node in the directed acyclic graph, wherein the dependency relationship includes the dependency resources of each node, and the dependency resources include a model; A second acquisition module 930 is used to acquire, during the execution of the target workflow, dependent resources of executed nodes in the directed acyclic graph as target dependent resources; A second determination module 940, configured to determine whether the target dependent resource is dependent on an unexecuted node of the directed acyclic graph based on the dependency relationship; The processing module 950 is used to release the target dependent resource when the target dependent resource is not dependent on the unexecuted node of the directed acyclic graph.

[0109] In an optional embodiment, the first determining module 920 is further configured to: Initialize a dependency map, where the dependency map is used to describe each of the dependent resources and the consumption nodes corresponding to the dependent resources; Traversing each node in the directed acyclic graph, and in the traversal process, determining a dependent resource of the current node as a first dependent resource based on an input of the current node; If the first dependent resource is not in the dependency map, constructing an output key of the first dependent resource and a consumer node set corresponding to the output key in the dependency map, and adding the current node to the consumer node set, wherein the output key includes a source node and an output index; In the case where the first dependent resource is already in the dependency map, the current node is added to a set of consumer nodes corresponding to the first dependent resource.

[0110] In an optional embodiment, the device 900 is further used for: When the first node is executed, traverse the set of consumer nodes corresponding to each output key in the dependency map, where the first node is any node of the directed acyclic graph; In the case where the first node is included in the consumer node set currently traversed, deleting the first node from the consumer node set currently traversed; When the currently traversed consumer node set is empty after the deletion, it is determined that the dependent resources corresponding to the currently traversed consumer node set are not dependent on the unexecuted nodes of the directed acyclic graph.

[0111] In an optional embodiment, the processing module 950 is further configured to: In a case where the target dependent resource is not dependent on an unexecuted node of the directed acyclic graph, determining whether a video memory pressure of the current GPU video memory is less than a first pressure threshold; When the video memory pressure is less than the first pressure threshold, the step of releasing the target dependent resource is skipped.

[0112] In an optional embodiment, the device 900 is further used for: When the video memory pressure is greater than or equal to the first pressure threshold, determining whether the video memory pressure is less than a second pressure threshold; When the video memory pressure is less than the second pressure threshold, transferring the target dependent resource from the GPU video memory to the CPU memory; When the video memory pressure is greater than or equal to the second pressure threshold, the target dependent resource is transferred from the GPU video memory to a disk.

[0113] In an optional embodiment, the device 900 is further used for: Before executing a second node of the directed acyclic graph, determining a dependent resource corresponding to the second node as a second dependent resource, wherein the second node is any node of the directed acyclic graph; Checking in sequence whether the second dependent resource is in the GPU video memory, the CPU memory or the disk; In the case that the second dependent resource is in the CPU memory or the disk, the second dependent resource is loaded into the GPU video memory.

[0114] In an optional embodiment, the device 900 is further used for: Intercepting target functions from third-party nodes, wherein the target functions include model loading, model unloading and migration operation functions; When the target function is intercepted, the workflow corresponding to the target function is added to the directed acyclic graph of the target workflow.

[0115] Embodiment 3 Fig.12 The schematic diagram of the hardware architecture of a computer device 10000 suitable for implementing a computer resource processing method according to the third embodiment of the present application is shown. In some embodiments, the computer device 10000 may be a terminal device such as a smart phone, a wearable device, a tablet computer, a personal computer, a vehicle terminal, a game console, a virtual device, a workbench, a digital assistant, a set-top box, a robot, etc. In other embodiments, the computer device 10000 may be a rack server, a blade server, a tower server, or a cabinet server (including an independent server, or a server cluster composed of multiple servers), etc. Fig.12 As shown, the computer device 10000 includes but is not limited to: a memory 10010, a processor 10020, and a network interface 10030 that can communicate with each other through a system bus. Among them: The memory 10010 includes at least one type of computer-readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 10010 can be an internal storage module of the computer device 10000, such as a hard disk or memory of the computer device 10000. In other embodiments, the memory 10010 can also be an external storage device of the computer device 10000, such as a plug-in hard disk equipped on the computer device 10000, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. Of course, the memory 10010 can also include both the internal storage module of the computer device 10000 and its external storage device. In this embodiment, the memory 10010 is generally used to store the operating system and various application software installed in the computer device 10000, such as program codes of computer resource processing methods, etc. In addition, the memory 10010 can also be used to temporarily store various data that have been output or are to be output.

[0116] In some embodiments, the processor 10020 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other chips. The processor 10020 is generally used to control the overall operation of the computer device 10000, such as performing control and processing related to data interaction or communication with the computer device 10000. In this embodiment, the processor 10020 is used to run the program code stored in the memory 10010 or process data.

[0117] The network interface 10030 may include a wireless network interface or a wired network interface, and the network interface 10030 is generally used to establish a communication link between the computer device 10000 and other computer devices. For example, the network interface 10030 is used to connect the computer device 10000 to an external terminal through a network, and to establish a data transmission channel and a communication link between the computer device 10000 and the external terminal. The network may be a wireless or wired network such as an intranet, the Internet, the Global System of Mobile communication (GSM), Wideband Code Division Multiple Access (WCDMA), 4G network, 5G network, Bluetooth, Wi-Fi, etc.

[0118] It should be pointed out that Fig.12 Only a computer device having components 10010 - 10030 is shown, but it should be understood that implementation of all of the components shown is not a requirement, and more or fewer components may alternatively be implemented.

[0119] In this embodiment, the computer resource processing method stored in the memory 10010 can also be divided into one or more program modules and executed by one or more processors (such as processor 10020) to complete the embodiment of the present application.

[0120] Embodiment 4 An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored, wherein when the computer program is executed by a processor, the steps of the computer resource processing method in the embodiment are implemented.

[0121] In this embodiment, the computer-readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (for example, SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, disk, optical disk, etc. In some embodiments, the computer-readable storage medium can be an internal storage unit of a computer device, such as a hard disk or memory of the computer device. In other embodiments, the computer-readable storage medium can also be an external storage device of a computer device, such as a plug-in hard disk equipped on the computer device, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. Of course, the computer-readable storage medium can also include both the internal storage unit of the computer device and its external storage device. In this embodiment, the computer-readable storage medium is generally used to store an operating system and various application software installed on the computer device, such as the program code of the computer resource processing method in the embodiment. In addition, the computer-readable storage medium can also be used to temporarily store various types of data that have been output or are to be output.

[0122] Embodiment 5 An embodiment of the present application also provides a computer program product, including a computer program, which implements the method in the above embodiment when executed by a processor.

[0123] Obviously, those skilled in the art should understand that the modules or steps of the above-mentioned embodiments of the present application can be implemented by general-purpose computer devices, they can be concentrated on a single computer device, or distributed on a network composed of multiple computer devices, optionally, they can be implemented by executable program codes of computer devices, so that they can be stored in a storage device and executed by the computer device, and in some cases, the steps shown or described can be executed in a different order from that herein, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.

[0124] It should be noted that the above are only preferred embodiments of the present application, and the patent protection scope of the present application is not limited thereto. Any equivalent structure or equivalent process transformation made using the contents of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A computer resource processing method, characterized in that: The method comprises: Get the directed acyclic graph of the target workflow; Determine the dependency relationship of each node in the directed acyclic graph, wherein the dependency relationship includes the dependent resources of each node, and the dependent resources include a model; During the execution of the target workflow, obtaining dependent resources of executed nodes in the directed acyclic graph as target dependent resources; Determining whether the target dependent resource is dependent on an unexecuted node of the directed acyclic graph based on the dependency relationship; When the target dependent resource is not dependent on any unexecuted node of the directed acyclic graph, the target dependent resource is released.

2. The method according to claim 1, characterized in that Determining the dependency relationship between nodes in the directed acyclic graph includes: Initialize a dependency map, where the dependency map is used to describe each of the dependent resources and the consumption nodes corresponding to the dependent resources; Traversing each node in the directed acyclic graph, and in the traversal process, determining a dependent resource of the current node as a first dependent resource based on an input of the current node; If the first dependent resource is not in the dependency map, constructing an output key of the first dependent resource and a consumer node set corresponding to the output key in the dependency map, and adding the current node to the consumer node set, wherein the output key includes a source node and an output index; In the case where the first dependent resource is already in the dependency map, the current node is added to a set of consumer nodes corresponding to the first dependent resource.

3. The method according to claim 2, characterized in that In the execution process of the target workflow, obtaining the dependent resources of the executed nodes in the directed acyclic graph as the target dependent resources, and determining whether the target dependent resources are dependent on the unexecuted nodes of the directed acyclic graph based on the dependency relationship, including: When the first node is executed, traverse the set of consumer nodes corresponding to each output key in the dependency map, where the first node is any node of the directed acyclic graph; In the case where the first node is included in the consumer node set currently traversed, deleting the first node from the consumer node set currently traversed; When the currently traversed consumer node set is empty after the deletion, it is determined that the dependent resources corresponding to the currently traversed consumer node set are not dependent on the unexecuted nodes of the directed acyclic graph.

4. The method according to claim 1, characterized in that: The step of releasing the target dependent resource when the target dependent resource is not dependent on any unexecuted node of the directed acyclic graph includes: In a case where the target dependent resource is not dependent on an unexecuted node of the directed acyclic graph, determining whether a video memory pressure of the current GPU video memory is less than a first pressure threshold; When the video memory pressure is less than the first pressure threshold, the step of releasing the target dependent resource is skipped.

5. The method according to claim 4, characterized in that The method further comprises: When the video memory pressure is greater than or equal to the first pressure threshold, determining whether the video memory pressure is less than a second pressure threshold; When the video memory pressure is less than the second pressure threshold, transferring the target dependent resource from the GPU video memory to the CPU memory; When the video memory pressure is greater than or equal to the second pressure threshold, the target dependent resource is transferred from the GPU video memory to a disk.

6. The method according to any one of claims 1 to 5, characterized in that: The method further comprises: Before executing a second node of the directed acyclic graph, determining a dependent resource corresponding to the second node as a second dependent resource, wherein the second node is any node of the directed acyclic graph; Checking in sequence whether the second dependent resource is in the GPU video memory, the CPU memory or the disk; In the case that the second dependent resource is in the CPU memory or the disk, the second dependent resource is loaded into the GPU video memory.

7. The method according to any one of claims 1 to 5, characterized in that: The method further comprises: Intercepting target functions from third-party nodes, wherein the target functions include model loading, model unloading and migration operation functions; When the target function is intercepted, the workflow corresponding to the target function is added to the directed acyclic graph of the target workflow.

8. A computer resource processing device, characterized in that: The device comprises: A first acquisition module is used to acquire a directed acyclic graph of a target workflow; A first determination module is used to determine the dependency relationship of each node in the directed acyclic graph, wherein the dependency relationship includes the dependent resources of each node, and the dependent resources include a model; A second acquisition module is used to acquire, during the execution of the target workflow, dependent resources of executed nodes in the directed acyclic graph as target dependent resources; A second determination module, configured to determine, based on the dependency relationship, whether the target dependent resource is dependent on an unexecuted node of the directed acyclic graph; The processing module is used to release the target dependent resource when the target dependent resource is not dependent on the unexecuted node of the directed acyclic graph.

9. A computer device, characterized in that: include: at least one processor; and a memory communicatively connected to the at least one processor; wherein: The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.

11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Task adjustment method applied to task engine, related device and storage medium

    CN113377348A

  • Dependent task unloading system and method based on mobile edge computing

    CN114980216A

  • Graph convolutional network enhanced edge computing task unloading method

    CN119847634A

  • Data processing and analyzing method and system of edge computing gateway

    CN119883661A

  • Application Management Platform for Hyper-Converged Cloud Infrastructures

    US20240411612A1

Cited By

  • SIMD architecture-oriented neural network processor intra-core scheduling method and system

    CN121833052A