Method for optimizing cache operation and computing equipment

By analyzing the resource and cache operation operation of the deep learning network model graph structure, and generating computing resource graphs and cache operation graphs, the problem of inaccurate cache operations in the deep learning network is solved, efficient and accurate cache operations are achieved, and graph execution efficiency is improved.

CN120104307APending Publication Date: 2025-06-06SHENZHEN CORERAIN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510042399.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In deep learning networks, it is difficult for the prior art to accurately identify which outputs or inputs need to be cached, resulting in unnecessary cache operations and affecting the efficiency of graph execution.

Method used

By loading the model diagram structure of the calculation model, perform resource analysis and cache operation analysis, generate calculation resource diagrams and cache operation diagrams, determine which outputs need to be cached, and optimize cache operations.

Benefits of technology

It realizes efficient and accurate caching operations, reduces unnecessary caching operations, improves graph execution efficiency, and improves system stability and reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104307A_ABST
    Figure CN120104307A_ABST
Patent Text Reader

Abstract

The invention provides a method for optimizing cache operation and computing equipment. The method comprises the following steps: loading a model graph structure of a computing model; performing resource analysis on the execution resource of the model graph structure to obtain a resource analysis result; performing cache operation analysis based on the resource analysis result to obtain a cache operation analysis result; and executing the model graph structure, and determining whether to perform cache operation on the output of the operator in the model graph structure according to the cache operation analysis result. According to the technical scheme, which outputs or inputs need to be subjected to cache operation can be accurately and accurately identified in advance, unnecessary cache operation is optimized, and the image execution efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of software development, and in particular to a method and computing device for optimizing cache operations. Background Art

[0002] In the deployment of deep learning networks, various hardware devices are usually selected for computing acceleration. When using computing acceleration hardware, not all operators can be executed on computing acceleration hardware. Model operators are divided into two categories: one is operators that can be executed on computing acceleration hardware, and the other is software operators executed by processors. When the above two types of operators are interspersed, in order to avoid too many additional memory copy operations, the software operators executed by the processor usually use the memory resources of computing acceleration hardware. The use of computing acceleration hardware memory resources by software operators involves cache operations.

[0003] In the actual computing process, not all processor operator outputs need to be updated to the memory of the computing acceleration hardware, and not all operator outputs executed by the computing acceleration hardware need to clear the cache of the operator output results.

[0004] To this end, a technical solution is needed that can accurately identify in advance which outputs or inputs need to be cached, optimize unnecessary cache operations, and improve the efficiency of graph execution. Summary of the invention

[0005] The present invention aims to provide a method and computing device for optimizing cache operations, which can better verify boundaries and special addresses, increase the completeness of chip verification, combine address space verification and functional verification, realize automatic address allocation, and increase verification efficiency.

[0006] According to one aspect of the present invention, a method for optimizing cache operation is provided, the method comprising:

[0007] Load the model graph structure of the computational model;

[0008] Performing resource analysis on the execution resources of the model graph structure to obtain resource analysis results;

[0009] Based on the resource analysis result, a cache operation analysis is performed to obtain a cache operation analysis result;

[0010] The model graph structure is executed, wherein whether to perform a cache operation on the output of the operator in the model graph structure is determined according to the cache operation analysis result.

[0011] According to some embodiments, loading a model graph structure of a computing model includes:

[0012] The topological information of the parsed calculation model is loaded into a model structure diagram, wherein the edges in the model structure diagram represent data transfer, and if the starting points of the arrows of the edges are the same, they represent the same data.

[0013] According to some embodiments, a resource analysis is performed on the execution resources of the model graph structure to obtain a resource analysis result, including: analyzing computing resources based on operator computing rules and the model graph structure to generate a computing resource graph.

[0014] According to some embodiments, based on the resource analysis result, cache operation analysis is performed to obtain a cache operation analysis result, including:

[0015] Based on the cache operation rules and the computing resource graph analysis, a cache operation graph is generated, and the cache operation graph defines the control flow of the execution graph.

[0016] According to some embodiments, determining whether to perform a cache operation on the output of the operator in the model graph structure according to the cache operation analysis result includes:

[0017] If the output of the software operator is used as the input of the hardware operator, the first cache operation is performed. The first cache operation is to update the execution result from the cache to the system memory after the execution of the software operator is completed.

[0018] According to some embodiments, determining whether to perform a cache operation on the output of the operator in the model graph structure according to the cache operation analysis result further includes:

[0019] If the output of the hardware operator is used as the input of the software operator, the second cache operation is performed, and the second cache operation is to clear the cache of the output result of the hardware operator after the computing acceleration hardware executes the hardware operator.

[0020] According to some embodiments, the control flow includes scheduling execution of all operators, input and output processing of all operators, and cache operations;

[0021] The cache operation graph maintains the cache operation rules of the model graph structure;

[0022] The mapping of all operator output names is used as the identifier for determining whether all operators perform cache operations.

[0023] According to some embodiments, determining whether to perform a cache operation on the output of the operator in the model graph structure according to the cache operation analysis result includes:

[0024] According to another aspect of the present invention, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the method as described above is implemented.

[0025] According to another aspect of the present invention, there is provided a computing device, comprising:

[0026] Processor; and

[0027] A memory stores a computer program, and when the computer program is executed by the processor, the method described in any one of the above items is implemented.

[0028] According to an embodiment of the present invention, the model graph structure of the computing model is first loaded, a resource analysis is performed on the execution resources of the model graph structure to obtain a resource analysis result, a cache operation analysis is performed based on the resource analysis result to obtain a cache operation analysis result, the model graph structure is executed, and it is determined whether to perform a cache operation on the output of the operator in the model graph structure according to the cache operation analysis result. The present invention determines which outputs require cache operations by performing resource analysis and cache operation analysis. When executing the graph, based on the model graph structure and cache operation analysis, cache operations can be performed efficiently and accurately to reduce unnecessary cache operations. By accurately identifying which outputs require cache operations in advance, unnecessary cache operations are optimized and the efficiency of graph execution is improved.

[0029] According to some embodiments, by judging computing resources and cache operations through a graph, it is possible to quickly identify which outputs need to be cached, and it is easier to plan and adjust cache strategies.

[0030] According to some embodiments, based on the analysis of cache operations, it is possible to more efficiently and accurately decide when and where to apply the cache during execution. This helps to reduce unnecessary cache updates or queries, thereby saving time and computing resources, ensuring that cache operations are performed as expected, and improving the stability and reliability of the system.

[0031] According to some embodiments, by analyzing computing resources, bottlenecks can be identified, and resource allocation can be optimized in a targeted manner. At the same time, reasonable use of cache can significantly reduce the pressure on backend services, speed up response time, and improve user experience.

[0032] According to some embodiments, when the system needs to be expanded or modified, having a clear resource analysis and cache operation analysis can make the process smoother and provide important information about the current system behavior. An effective cache strategy can directly affect the cost of cloud services or other hosting solutions. By precisely controlling cache operations, expenses can be minimized without affecting performance.

[0033] It is to be understood that the foregoing general description and the following detailed description are exemplary only and are not restrictive of the invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for describing the embodiments are briefly introduced below.

[0035] Figure 1 A flow chart of a method for optimizing cache operations according to an example embodiment is shown.

[0036] Figure 2 A schematic diagram illustrating a model graph structure according to an example embodiment.

[0037] Figure 3 A schematic diagram illustrating a computing resource graph according to an example embodiment.

[0038] Figure 4 A schematic diagram illustrating a cache operation graph according to an example embodiment.

[0039] Figure 5 A block diagram of a computing device is shown according to an exemplary embodiment. DETAILED DESCRIPTION

[0040] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in many forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that the present invention will be comprehensive and complete and fully convey the concepts of the example embodiments to those skilled in the art. The same reference numerals in the figures represent the same or similar parts, and thus their repeated description will be omitted.

[0041] In addition, the described features, structures or characteristics may be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present invention. However, those skilled in the art will appreciate that the technical solution of the present invention can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be adopted. In other cases, known methods, devices, implementations or operations are not shown or described in detail to avoid blurring various aspects of the present invention.

[0042] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0043] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to actual conditions.

[0044] It should be understood that although the terms first, second, third, etc. may be used herein to describe various components, these components should not be limited by these terms. These terms are used to distinguish one component from another component. Therefore, the first component discussed below can be referred to as the second component without departing from the teachings of the present inventive concept. As used herein, the term "and / or" includes any one of the associated listed items and all combinations of one or more.

[0045] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the present invention are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0046] Those skilled in the art will appreciate that the drawings are merely schematic diagrams of example embodiments, and the modules or processes in the drawings are not necessarily necessary for implementing the present invention, and therefore cannot be used to limit the protection scope of the present invention.

[0047] In order to ensure the correctness of the calculation results, usually after the execution of the software operator is completed, its results need to be updated from the cache to the memory of the computing acceleration hardware, and after the computing acceleration hardware executes the hardware operator, the cache of the operator output results needs to be cleared. In the actual process, not all software operator outputs need to update their results from the cache to the memory of the computing acceleration hardware, and not all hardware operator outputs executed by the computing acceleration hardware need to clear the cache of the operator output results. For example, when the output of a software operator is only the input of other software operators, the output does not need to be updated to the memory of the computing acceleration hardware. When the output of a hardware operator executed by the computing acceleration hardware is only the input of a hardware operator executed by other computing acceleration hardware, the output does not need to clear the cache of the operator output results. The output of the software operator needs to be given to the computing acceleration hardware operator as input, the output needs to be updated to the memory of the computing acceleration hardware. The output of the hardware operator executed by the computing acceleration hardware needs to be given to the software operator as input, the cache of the operator output results needs to be cleared.

[0048] There are two main ways of traditional cache operation schemes. One is to update the output of all software operators from the cache to the memory of the computing acceleration hardware, and to clear the result cache of all computing acceleration hardware operator outputs; the other is that the software operators do not use the memory of the computing acceleration hardware, and add memory copy operations between the computing acceleration hardware operators and the software operators. The first method performs many unnecessary cache operations, and the second method performs many more memory copy operations.

[0049] To this end, the present invention proposes a method for optimizing cache operations, which can better verify boundaries and special addresses, increase the completeness of chip verification, combine address space verification and functional verification, realize automatic address allocation, and increase verification efficiency.

[0050] Exemplary embodiments of the present invention are described below with reference to the accompanying drawings.

[0051] Figure 1 A flow chart of a method for optimizing cache operations according to an example embodiment is shown.

[0052] See also Figure 1 , in S101, the model graph structure of the calculation model is loaded.

[0053] According to some embodiments, the topological information of the parsed computing model is loaded into a model structure diagram; the edges in the model structure diagram represent data transfer, and if the starting points of the arrows of the edges are the same, they represent the same data.

[0054] According to some embodiments, a process control diagram is generated based on the rules of an artificial intelligence chip (AI chip). Operators that cannot be executed by hardware on the chip are transferred to the processor for execution, that is, the processor executes software operators (CPU OP). Software operators are not independent of the model, but are also part of the model. The model is loaded, the topological information of the model is parsed, and the model graph structure is constructed, where the edges in the graph are data transfers, and the arrows with the same starting point are the same data. The graph structure fragment is as follows: Figure 2 The schematic diagram of the model graph structure is shown in the figure.

[0055] In S103, a resource analysis is performed on the execution resources of the model graph structure to obtain a resource analysis result.

[0056] According to some embodiments, computing resources are analyzed based on operator computing rules and the model graph structure to generate a computing resource graph. According to the model graph structure obtained in S101, based on the information that the computing acceleration hardware has its corresponding operator support list and the constraints of the supported operators, the hardware operators (hardware OP) supported by the computing acceleration hardware are deployed on the computing acceleration hardware, and the operators not supported by the computing acceleration hardware are deployed on the processor. The computing resource graph is generated based on the above rules, such as Figure 3 A schematic diagram of the computing resource graph is shown.

[0057] In S105, cache operation analysis is performed based on the resource analysis result to obtain a cache operation analysis result.

[0058] According to some embodiments, a cache operation graph is generated based on the cache operation rules and the computing resource graph analysis; the cache operation graph defines the control flow of the execution graph. The control flow includes the scheduling execution of all operators, the input and output processing of all operators, and the cache operation; the cache operation graph maintains the cache operation rules of the model graph structure; the mapping of all operator output names is used as an identifier to determine whether all operators perform cache operations.

[0059] According to some embodiments, the cache operation graph defines the data mapping relationship between cache identifiers, cache attributes, attribute variables, and data structures. The processor executes the control flow of the cache operation graph; the control flow includes the scheduling execution of all operators, the input and output processing of all operators, and the cache operation. The cache operation graph maintains the rules of the cache operation of the model graph structure; the mapping of all operator output names is used as an identifier to determine whether all operators perform cache operations. When all operators are executed, the corresponding cache identifier is obtained from the cache operation graph through the output operator name to confirm whether to perform the cache operation.

[0060] In order to ensure the correct execution of the results, the outputs transmitted between the hardware operator and the software operator need to be cached (cache operation). The first cache operation is to update the execution result from the cache to the system memory after the software operator is executed, which is called flush; the second cache operation is to clear the cache of the hardware operator output result after the computing acceleration hardware executes the hardware operator, which is called invalid. The cache operation graph actually generates cache identification information based on the original graph and saves the cache operation graph. The generation rules are: determine whether the CPU OP output is given to the computing acceleration hardware OP as input. If not, no flush is required. If yes, flush is required; determine whether the output of the computing acceleration hardware OP is given to the CPU OP as input. If not, no invalidation is required. If yes, invalidation is required. In the cache operation graph, the edges connecting the two operators require cache operations, otherwise the results will be wrong or random.

[0061] Based on this rule and graph analysis, we can get the cache operation graph, such as Figure 4 The schematic diagram of the cache operation graph is shown below. The cache operation graph is an independent graph that does not change the original graph and is constructed during the loading phase of the original graph. The cache operation graph maintains the cache operation rules of the original graph and is mapped by the names of all operator outputs. The cache operation graph defines some cache identifiers, attributes, attribute variables, etc. that indicate whether to perform cache operations, that is, to represent the data structure with a graph and maintain a mapping table.

[0062] According to some embodiments, the CPU executes a control flow in a cache operation graph, where the control flow includes operator scheduling execution, OP input and output processing, cache operations, and the like.

[0063] In S107, the model graph structure is executed, wherein it is determined whether to perform a cache operation on the output of the operator in the model graph structure according to the cache operation analysis result.

[0064] According to some embodiments, according to the cache operation analysis result, if the output of the software operator is used as the input of the hardware operator, the first cache operation is performed; according to the cache operation analysis result, if the output of the hardware operator is used as the input of the software operator, the second cache operation is performed. The first cache operation is to update the execution result from the cache to the system memory after the software operator is executed; the second cache operation is to clear the cache of the hardware operator output result after the computing acceleration hardware executes the hardware operator.

[0065] According to some embodiments, the cache operation graph obtained in S105 is combined with the model graph structure obtained in S101 to start executing the cache operation graph. When all operators are executed according to the model graph structure, the corresponding cache identifier is obtained from the cache operation graph through the output name of the operator output to confirm whether to perform the cache operation. When the cache operation graph is actually executed, it is only necessary to perform cache operations on all operator outputs according to the cache operation graph. When the operator execution is completed, the corresponding cache identifier is obtained from the cache operation graph through the name of the operator output to confirm whether to perform the cache operation.

[0066] The present invention first loads the model graph structure of the computing model, generates a computing resource graph according to the model graph structure, generates a cache operation graph according to the computing resource graph, and finally executes the cache operation graph by combining the cache operation graph with the model graph structure. The present invention can determine which outputs need to be cached by generating a computing resource graph and a cache operation graph. When executing the graph, based on the cache operation graph, cache operations can be performed efficiently and accurately, unnecessary cache operations can be reduced, and which outputs need to be cached can be accurately and precisely identified in advance, unnecessary cache operations can be optimized, and the efficiency of graph execution can be improved.

[0067] All graph constructions of the present invention are fully automatically executed. The graph loading process automatically analyzes the graph and automatically generates cache identifiers on the nodes on the graph to optimize cache operations, so as to minimize the number of cache operations and optimize graph execution. By generating a computing resource graph and a cache operation graph, it is possible to determine which outputs need to be cached. When executing the graph, based on the cache operation graph, cache operations can be performed efficiently and accurately, reducing unnecessary cache operations.

[0068] Figure 5 A block diagram of a computing device is shown according to an exemplary embodiment.

[0069] like Figure 5 As shown, computing device 30 includes processor 12 and memory 14. Computing device 30 may also include bus 22, network interface 16, and I / O interface 18. Processor 12, memory 14, network interface 16, and I / O interface 18 may communicate with each other via bus 22.

[0070] The processor 12 may include one or more general-purpose CPUs (Central Processing Units, processors), microprocessors, or application-specific integrated circuits, etc., for executing relevant program instructions. According to some embodiments, the computing device 30 may also include a high-performance graphics card (GPU) 20 for accelerating the processor 12.

[0071] The memory 14 may include a machine system readable medium in the form of a volatile memory, such as a random access memory (RAM), a read-only memory (ROM) and / or a cache memory. The memory 14 is used to store one or more programs including instructions and data. The processor 12 can read the instructions stored in the memory 14 to execute the above-mentioned method according to the embodiment of the present invention.

[0072] The computing device 30 may also communicate with one or more networks via the network interface 16. The network interface 16 may be a wireless network interface.

[0073] The bus 22 may include an address bus, a data bus, a control bus, etc. The bus 22 provides a path for exchanging information between components.

[0074] It should be noted that, in the specific implementation process, the computing device 30 may also include other components necessary for normal operation. In addition, those skilled in the art may understand that the above device may only include components necessary for implementing the embodiments of this specification, and need not include all components shown in the figure.

[0075] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of the above method when executed by a processor. The computer-readable storage medium may include, but is not limited to, any type of disk, including a floppy disk, an optical disk, a DVD, a CD-ROM, a microdrive, and a magneto-optical disk, a ROM, a RAM, an EPROM, an EEPROM, a DRAM, a VRAM, a flash memory device, a magnetic card or an optical card, a nanosystem (including a molecular memory IC), a network storage device, a cloud storage device, or any type of medium or device suitable for storing instructions and / or data.

[0076] An embodiment of the present invention further provides a computer program product, which includes a computer program. The computer program is operable to enable a computer to execute part or all of the steps of any one of the methods described in the above method embodiments.

[0077] Those skilled in the art can clearly understand that the technical solution of the present invention can be implemented with the help of software and / or hardware. "Unit" and "module" in this specification refer to software and / or hardware that can independently complete or cooperate with other components to complete specific functions, where the hardware can be, for example, a field programmable gate array, an integrated circuit, etc.

[0078] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present invention.

[0079] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0080] In the several embodiments provided by the present invention, it should be understood that the disclosed devices can be implemented in other ways. For example, the device embodiments described above are only schematic, such as the division of units, which is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some service interfaces, and the indirect coupling or communication connection of devices or units can be electrical or other forms.

[0081] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0082] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0083] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or the whole or part of the technical solution can be embodied in the form of a software product, which is stored in a memory and includes several instructions for a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the methods of various embodiments of the present invention.

[0084] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0085] The exemplary embodiments of the present invention are specifically shown and described above. It should be understood that the present invention is not limited to the detailed structures, configurations or implementations described herein; on the contrary, the present invention is intended to cover various modifications and equivalent configurations included in the spirit and scope of the appended clauses.

Claims

1. A method for optimizing cache operation, the method comprising: Load the model graph structure of the computational model; Performing resource analysis on the execution resources of the model graph structure to obtain resource analysis results; Based on the resource analysis result, a cache operation analysis is performed to obtain a cache operation analysis result; The model graph structure is executed by a processor and acceleration hardware, wherein whether to perform a cache operation on the output of the operator in the model graph structure is determined according to the cache operation analysis result.

2. The method according to claim 1, characterized in that: Load the model graph structure of the calculation model, including: The topological information of the parsed calculation model is loaded into a model structure diagram, wherein the edges in the model structure diagram represent data transfer, and if the starting points of the arrows of the edges are the same, they represent the same data.

3. The method according to claim 1, characterized in that: Performing resource analysis on the execution resources of the model graph structure to obtain resource analysis results, including: analyzing computing resources based on operator computing rules and the model graph structure to generate a computing resource graph.

4. The method according to claim 1, characterized in that Based on the resource analysis result, cache operation analysis is performed to obtain cache operation analysis results, including: Based on the cache operation rules and the computing resource graph analysis, a cache operation graph is generated, and the cache operation graph defines the control flow of the execution graph.

5. The method according to claim 1, characterized in that: Determining whether to perform a cache operation on the output of the operator in the model graph structure according to the cache operation analysis result includes: If the output of the software operator is used as the input of the hardware operator, the first cache operation is performed. The first cache operation is to update the execution result from the cache to the system memory after the execution of the software operator is completed.

6. The method according to claim 1, characterized in that Determining whether to perform a cache operation on the output of the operator in the model graph structure according to the cache operation analysis result, further comprising: If the output of the hardware operator is used as the input of the software operator, the second cache operation is performed, and the second cache operation is to clear the cache of the output result of the hardware operator after the computing acceleration hardware executes the hardware operator.

7. The method according to claim 4, characterized in that The control process includes the scheduling execution of all operators, the input and output processing of all operators, and the cache operation; The cache operation graph maintains the cache operation rules of the model graph structure; The mapping of all operator output names is used as the identifier for determining whether all operators perform cache operations.

8. The method according to claim 1, characterized in that Determining whether to perform a cache operation on the output of the operator in the model graph structure according to the cache operation analysis result includes: When all operators are executed according to the model graph structure, the corresponding cache identifier is obtained from the cache operation graph through the output name of the operator to confirm whether to perform the cache operation.

9. A computer program product, characterized in that The method comprises a computer program, which implements the method according to any one of claims 1 to 8 when being executed by a processor.

10. A computing device, characterized in that include: processor; as well as A memory storing a computer program, wherein when the computer program is executed by the processor, the method according to any one of claims 1 to 8 is implemented.