Processing method and device for loop code, electronic equipment and storage medium

By storing the target variables in the loop code to the global storage unit and executing code snippets using the cache space of the local storage unit, the storage overflow problem of the artificial intelligence chip when executing the loop code is solved, and memory access efficiency is improved.

CN119938004AActive Publication Date: 2025-05-06KUNLUNXIN TECHNOLOGY (BEIJING) CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510087487.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-06
Estimated Expiration
2045-01-20

AI Technical Summary

Technical Problem

When the artificial intelligence chip executes cyclic code, due to insufficient storage space of local storage units, data overflow, which increases memory access delay and reduces memory access efficiency.

Method used

Reduce memory access overhead by storing the target variables in multiple variables to the global storage unit and storing and executing the target variables of the code snippet in the local storage unit using the free cache space.

Benefits of technology

It effectively reduces the memory access overhead of target variables during loop code execution, improves memory access efficiency, and reduces delays due to data overflow.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938004A_ABST
    Figure CN119938004A_ABST
Patent Text Reader

Abstract

The invention provides a processing method for loop codes, and relates to the technical field of artificial intelligence, in particular to the technical field of chips and the technical field of storage. According to the specific implementation scheme, in response to determining that the total data volume of multiple variables of a loop code block is larger than the available storage space capacity of a local storage unit, multiple target variables in the multiple variables are stored in a global storage unit, and the loop code block comprises multiple code snippets; in response to determining that a first code snippet in the multiple code snippets is executed, a first target variable of the first code snippet in the multiple target variables is stored in a global storage unit and an idle first cache space exists in a local storage unit, storing the first target variable of the first code snippet in the idle first cache space, and executing the first code snippet based on the first target variable in the first cache space in the execution process of the loop code block. The invention further provides a code circulation device, electronic equipment and a storage medium.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to the field of chip technology and storage technology. More specifically, the present disclosure provides a processing method, device, electronic device and storage medium for loop code. Background Art

[0002] With the development of artificial intelligence technology, the application of artificial intelligence chips is increasing. Artificial intelligence chips can include local memory units. The core of the artificial intelligence chip can quickly access the data in the local memory unit. Summary of the invention

[0003] The present disclosure provides a method, apparatus, device and storage medium for processing loop codes.

[0004] According to one aspect of the present disclosure, a processing method for loop code is provided, the method comprising: in response to determining that the total data amount of multiple variables of a loop code block is greater than the available storage space capacity of a local storage unit, storing multiple target variables among the multiple variables in a global storage unit, wherein the loop code block includes multiple code fragments and the local storage unit includes at least one first cache space; in response to determining that a first code fragment among the multiple code fragments is executed, a first target variable of the first code fragment among the multiple target variables is stored in the global storage unit and determining that there is an idle first cache space in the local storage unit, storing the first target variable of the first code fragment in the idle first cache space, so as to execute the first code fragment based on the first target variable in the first cache space during the execution of the loop code block.

[0005] According to another aspect of the present disclosure, a processing device for loop code is provided, the device comprising: a first storage module, for storing multiple target variables among the multiple variables in a global storage unit in response to determining that the total data amount of multiple variables of the loop code block is greater than the available storage space capacity of a local storage unit, wherein the loop code block includes multiple code fragments and the local storage unit includes at least one first cache space; a first execution module, for storing the first target variable of the first code fragment in the free first cache space in response to determining that a first code fragment among the multiple code fragments is executed, a first target variable of the first code fragment among the multiple target variables is stored in the global storage unit and determining that there is a free first cache space in the local storage unit, so as to execute the first code fragment based on the first target variable in the first cache space during the execution of the loop code block.

[0006] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method provided according to the present disclosure.

[0007] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided. The computer instructions are used to cause a computer to execute the method provided according to the present disclosure.

[0008] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the method provided according to the present disclosure is implemented.

[0009] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.

[0011] Figure 1 is a flowchart of a method for processing loop code according to one embodiment of the present disclosure;

[0012] FIG. 2A to FIG. 2D is a schematic diagram of an artificial intelligence chip according to an embodiment of the present disclosure;

[0013] Figure 3 is a block diagram of a processing device for loop code according to one embodiment of the present disclosure; and

[0014] Figure 4 is a block diagram of an electronic device to which a processing method for loop codes can be applied according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0015] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0016] The capacity of the local storage unit is small, which is smaller than the capacity of the global memory unit. The program executed by the artificial intelligence chip can be called a kernel function. The local variables in the kernel function will occupy a large amount of storage space in the local storage unit. If the storage space of the local storage unit occupied by multiple local variables is greater than the available storage space capacity of the local storage unit, there will be a local storage overflow problem. The overflowed data can be stored in the global storage unit. However, the processing unit of the artificial intelligence chip cannot directly access the global storage unit. The processing unit of the artificial intelligence chip can be the above-mentioned core. If the overflowed data needs to be used, the overflowed data can be stored in the storage space of the local storage unit except the available storage space. During the code loop, one value of the array can be used in each loop. If the data is overflowed data, in order to use the data of the array, multiple memory accesses will occur between the global storage unit and the local storage unit, resulting in a high memory access delay, which in turn leads to low memory access efficiency.

[0017] Therefore, in order to fully improve the memory access efficiency, the present disclosure provides a processing method for loop code, which will be described below.

[0018] Figure 1 is a flowchart of a method for processing loop codes according to an embodiment of the present disclosure.

[0019] like Figure 1 As shown, the method 100 may include operations S110 to S120.

[0020] In operation S110 , in response to determining that a total data amount of a plurality of variables of a loop code block is greater than an available storage space capacity of a local storage unit, a plurality of target variables among the plurality of variables are stored in a global storage unit.

[0021] In the embodiment of the present disclosure, the plurality of variables may include complex variables and simple variables. The plurality of simple variables may include at least one of the following types of variables: integer variables, floating point variables, etc. The complex variable is, for example, an array variable or a structure variable.

[0022] In the disclosed embodiment, the kernel function may include one or more loop code blocks. The loop code block may include multiple code snippets. For example, in each loop process, one or more code snippets in the multiple code snippets may use the target variable.

[0023] In the embodiment of the present disclosure, the target variable may include one or more data. The target variable may be a complex variable.

[0024] In the disclosed embodiment, the capacity of the available storage space may be smaller than the capacity of the local storage unit. In the case where the total data volume of multiple target variables of the loop code block is larger than the capacity of the available storage space, the multiple target variables cannot all be stored in the local storage unit. The multiple target variables may all be stored in the global storage unit.

[0025] In the embodiment of the present disclosure, the local storage unit includes at least one first cache space. For example, one or more storage spaces in the local storage unit can be used as one or more first cache spaces.

[0026] In operation S120, in response to determining that a first code snippet among multiple code snippets is executed, a first target variable of the first code snippet among multiple variables is stored in a global storage unit, and determining that there is a free first cache space in a local storage unit, the first target variable of the first code snippet is stored in the free first cache space, so that the first code snippet is executed based on the first target variable in the first cache space during the execution of the loop code block.

[0027] In the embodiment of the present disclosure, the first code snippet may be determined from a plurality of code snippets. For example, one or more first code snippets may be determined based on the number of first cache spaces. The number of first cache spaces may be consistent with the number of first code snippets.

[0028] In the disclosed embodiment, after the first target variable of the first code snippet is stored in the first cache space and before the loop code block is executed, the first target variable of the first code snippet can be continuously stored in the first cache space. After the loop code block is executed, the target variable of the first code snippet or the processed target variable can be written to the global storage unit.

[0029] Through the disclosed embodiment, before the loop code block is executed, all complex variables are stored in the global storage unit. During the execution of the loop code block, the target variable of the first code snippet of the loop code block is always stored in the local storage unit, reducing the memory access overhead of the target variable of the first code snippet. The memory access overhead can be the overhead of moving the target variable of the first code snippet between the local storage unit and the global storage unit. The memory access delay can be effectively reduced and the memory access efficiency can be improved.

[0030] It can be understood that the method of the present disclosure is described above. FIG. 2A to FIG. 2D The method of the present disclosure is further described.

[0031] FIG. 2A to FIG. 2D is a schematic diagram of an artificial intelligence chip according to an embodiment of the present disclosure.

[0032] like FIG. 2A to FIG. 2DAs shown, the chip 200 may include a local storage unit lm20 and a plurality of processing units. The plurality of processing units may include a processing unit PU201, a processing unit PU202, and a processing unit PU203. The local storage unit lm20 may include a first cache space lm211, a first cache space lm212, and a first cache space lm213. The local storage unit lm20 may also include a second cache space lm221.

[0033] As shown in FIG. 2 , the chip 200 can read data from the global memory unit gm20 , and can also write data to the global memory unit gm20 .

[0034] In some embodiments, in some implementations of the above operation S110, in response to determining that the total data volume of multiple variables of the loop code block is greater than the available storage space capacity of the local storage unit, multiple target variables among the multiple variables are stored in the global storage unit. As shown in Figure 2, the total capacity of the first cache space lm211, the first cache space lm212, and the first cache space lm213 can be used as the available storage space capacity. If the total data volume of multiple variables of the loop code block is greater than the available storage space capacity, the overflow risk of the local storage unit is relatively large, and the multiple target variables can be stored in the global storage unit first.

[0035] In some embodiments, the multiple target variables include at least one of the following types of variables: array variables and structure variables. For example, taking the target variable as an array variable as an example, the multiple target variables may include array variable array21, array variable array22, array variable array23, array variable array24, and array variable array25. It is understood that the target variable may also be a structure variable, or the multiple target variables may include at least one array variable and at least one structure variable.

[0036] In some embodiments, storing the plurality of target variables in the global storage unit may include: storing the plurality of target variables in a plurality of continuous storage spaces of the global storage unit. Figure 2AAs shown, the global storage unit gm20 may include a plurality of continuous storage spaces. The plurality of continuous storage spaces may include storage space gm201, storage space gm202, storage space gm203, storage space gm204, and storage space gm205. Array variable array21 may be stored in storage space gm201. Array variable array22 may be stored in storage space gm202. Array variable array23 may be stored in storage space gm203. Array variable array24 may be stored in storage space gm204. Array variable array25 may be stored in storage space gm205. Through the embodiments of the present disclosure, multiple target variables are stored in multiple storage spaces that are read continuously, and data required for calculation may be quickly located based on offset values, thereby improving memory access efficiency.

[0037] In some embodiments, in some implementations of the above operation S120, in response to determining that a first code snippet among multiple code snippets is executed, a target variable of the first code snippet is stored in a global storage unit, and determining that there is a free first cache space in a local storage unit, the first target variable of the first code snippet is stored in the free first cache space.

[0038] For example, in the first loop of the loop code block, the first three code snippets of the loop code block can be used as three first code snippets. When executing the first code snippet, the first target variable of the first code snippet can be an array variable array21. The array variable array21 is stored in the global storage unit. The local storage unit lm20 includes three free first cache spaces. Figure 2B As shown, the array variable array21 can be loaded from the storage space gm201 to the first cache space lm211.

[0039] In some embodiments, executing the first code snippet based on the first target variable in the first cache space during the execution of the loop code block may include: obtaining at least one data of the target variable of the first code snippet from the first cache space. For example, the processing unit PU201 may obtain one or more data of the array variable array21 from the first cache space lm211 to perform calculations and obtain an intermediate result of the first code snippet. Through the embodiment of the present disclosure, during the execution of the loop code, the first target variable is stored in the first cache space of the local storage unit unit, and data can be efficiently obtained from the local storage unit to improve memory access efficiency. Compared with the local storage unit that is not set to the first cache space, the memory access overhead required for repeatedly loading the first target variable in different loop rounds is saved, and the execution efficiency of the loop code can be fully improved. For artificial intelligence chips that do not have a first-level cache (L1 cache) or a second-level cache (L2 cache), the memory access overhead can be reduced more efficiently.

[0040] In some embodiments, executing the first code snippet based on the first target variable in the first cache space during the execution of the loop code block may also include: using at least one intermediate result of the loop code block to replace at least one data of the first target variable of the first code snippet in the first cache space. For example, an intermediate result of the first code snippet can be used to replace a data in the array variable array21. Through the embodiment of the present disclosure, during the execution of the loop code, the first target variable is stored in the first cache space of the local storage unit unit, and data can be efficiently written to the array variable, effectively improving the memory access efficiency. Compared with the local storage unit that is not set to the first cache space, the memory access overhead required for repeatedly loading the first target variable in different loop rounds is saved, and the execution efficiency of the loop code can be fully improved. For artificial intelligence chips that do not have a first-level cache (L1 cache) or a second-level cache (L2 cache), the memory access overhead can be reduced more efficiently.

[0041] For another example, when executing the second code snippet, the target variable of the second code snippet may be array variable array22. Array variable array22 is stored in the global storage unit. Local storage unit lm20 includes two free first cache spaces. Array variable array22 may be loaded from storage space gm202 to first cache space lm212. Next, data of array variable array22 may be obtained from first cache space lm212. When executing the third code snippet, the target variable of the third code snippet may be array variable array23. Array variable array23 is stored in the global storage unit. Local storage unit lm20 includes one free first cache space. Array variable array23 may be loaded from storage space gm203 to first cache space lm213. Next, data of array variable array23 may be obtained from first cache space lm213.

[0042] In some embodiments, in response to determining that a second code snippet among multiple code snippets is executed, a target variable of the second code snippet is stored in a global storage unit, and determining that there is no free first cache space in a local storage unit, the second code snippet is executed based on the second target variable in the global storage unit during the execution of the loop code block.

[0043] For example, during the first loop of the loop code block, the 4th to 5th code snippets of the loop code block can be used as two second code snippets. When executing the 4th code snippet, the target variable of the 4th code snippet can be an array variable array24. The array variable array24 is stored in the global storage unit. The three first cache spaces of the local storage unit lm20 are all occupied. The array variable array24 can be loaded from the storage space gm204 to the second cache space lm221.

[0044] In some embodiments, executing the second code snippet based on the second target variable in the global storage unit during the execution of the loop code block includes: obtaining at least one data of the second target variable of the second code snippet from the second cache space. For example, when executing the fourth code snippet, the data of the array variable array24 can be obtained from the second cache space lm221 to perform calculations to obtain an intermediate result of the fourth code snippet. Through the embodiments of the present disclosure, a second cache space is set to cache the second target variable to avoid the second target variable occupying the first cache space, and the first target variable can be continuously stored in the local storage space during the execution of the loop code block, which can effectively save memory access overhead and improve overall memory access efficiency.

[0045] In some embodiments, executing the second code snippet based on the second target variable in the global storage unit during the execution of the loop code block includes: replacing at least one data of the second target variable of the second code snippet in the global storage unit with at least one intermediate result of the loop code block. For example, the data in the array variable array24 can be replaced with the intermediate result to obtain the array variable array24_1.

[0046] In some embodiments, executing the second code segment based on the second target variable in the global storage unit during the execution of the loop code block includes: in response to determining that the second code segment is executed and determining that at least one second cache space is occupied, releasing at least one occupied second cache space to obtain a released second cache space. Figure 2C As shown, during the first loop of the loop code block, when the fifth code snippet is executed, the target variable of the fifth code snippet can be the array variable array25. The array variable array25 is stored in the global storage unit. The three first cache spaces of the local storage unit lm20 are all occupied. The second cache space lm221 is also occupied. The array variable array24_1 can be loaded from the second cache space lm221 to the storage space gm204 to release the second cache space lm221. As mentioned above, compared with the array variable array24, a data in the array variable array24_1 is replaced with the intermediate result of the fourth code snippet. After the array variable array24 after the first replacement is loaded into the storage space gm204, the replacement of the data in the second target variable of the second code snippet in the global storage unit can be realized.

[0047] In some embodiments, executing the second code snippet based on the second target variable in the global storage unit during the execution of the loop code block includes: loading the target variable of the second code snippet from the global storage unit to the released second cache space, so as to obtain at least one data of the target variable of the second code snippet from the released second cache space. Figure 2D As shown, after releasing the second cache space lm221, the array variable array25 can be loaded from the storage space gm205 to the second cache space lm221. Next, when executing the fifth code snippet, at least one data of the array variable array25 can be obtained from the second cache space lm221.

[0048] For example, after completing the first loop of the loop code block, the second loop can be executed. During the second loop, when executing the first code snippet, the data of the array variable can be obtained from the first cache space lm211. When executing the second code snippet, the data of the array variable can be obtained from the first cache space lm212. When executing the third code snippet, the data of the array variable can be obtained from the first cache space lm212.

[0049] For another example, during the execution of the second loop and when executing the fourth code snippet, the array variable array25 can be moved from the second cache space lm221 to the storage space gm205 to release the second cache space lm221. Next, the array variable array24_1 can be loaded from the storage space gm204 to the second cache space lm221. During the execution of the second loop and when executing the fourth code snippet, the data of the array variable array24_1 can be obtained from the second cache space lm221 to perform calculations and obtain another intermediate result of the fourth code snippet. The intermediate result can be used to replace the data in the array variable array24_1 to obtain the array variable after two replacements.

[0050] Next, the manner of executing the fifth code snippet of the second loop is the same as or similar to the manner of executing the fifth code snippet of the first loop, and the present disclosure will not elaborate on it here.

[0051] It can be understood that the method of executing the third cycle to the last cycle is the same or similar to the method of executing the second cycle, and the present disclosure will not repeat them here.

[0052] In some embodiments, the method may further include: in response to determining that the loop code block is executed, writing the processed target variable of the first code fragment into a global storage unit. For example, after the loop end condition is met, the three array variables in the first cache space lm211, the first cache space lm212, and the first cache space lm213 can be written into the storage space gm201, the storage space gm202, and the storage space gm203 respectively as processed array variables. The loop end condition may be: the number of loops is equal to a preset number threshold. Through the embodiment of the present disclosure, during the execution of the loop code block, the first target variable can be stored in a local storage unit. After the loop code block is executed, the first target variable is written into the global storage unit. The memory access overhead can be effectively reduced and the memory access efficiency can be improved.

[0053] It can be understood that the method of the present disclosure is described above, and the device of the present disclosure will be described below.

[0054] Figure 3is a block diagram of a processing apparatus for loop code according to an embodiment of the present disclosure.

[0055] like Figure 3 As shown, the device 300 may include a first processing module 310 and a first execution module 320 .

[0056] The first storage module 310 is used to store multiple target variables among the multiple variables in the global storage unit in response to determining that the total data amount of the multiple variables of the loop code block is greater than the available storage space capacity of the local storage unit. The loop code block includes multiple code fragments, and the local storage unit includes at least one first cache space.

[0057] The first execution module 320 is used to store the first target variable of the first code snippet to the free first cache space in response to determining that a first code snippet among multiple code snippets is executed, a first target variable of the first code snippet among multiple target variables is stored in a global storage unit, and determining that there is a free first cache space in a local storage unit, so as to execute the first code snippet based on the first target variable in the first cache space during the execution of the loop code block.

[0058] In some embodiments, the device 300 also includes: a second execution module, which is used to execute the second code snippet based on the second target variable in the global storage unit during the execution of the loop code block in response to determining that the second code snippet among multiple code snippets is executed, the second target variable of the second code snippet among multiple target variables is stored in the global storage unit, and determining that there is no free first cache space in the local storage unit.

[0059] In some embodiments, the first processing module includes: a first processing submodule, configured to store the plurality of target variables into a plurality of continuous storage spaces of the global storage unit.

[0060] In some embodiments, the plurality of target variables include at least one of the following types of variables: an array variable and a structure variable.

[0061] In some embodiments, the first execution module includes at least one of the following operations: an acquisition submodule, configured to acquire at least one data of the first target variable of the first code snippet from the first cache space. A first replacement submodule, configured to replace at least one data of the first target variable of the first code snippet in the first cache space with at least one intermediate result of the loop code block.

[0062] In some embodiments, the apparatus 300 further includes: a writing module, configured to write the processed target variable of the first code segment into a global storage unit in response to determining that the execution of the loop code block is completed.

[0063] In some embodiments, the local storage unit further includes at least one second cache space. The second execution module includes: a release submodule, for releasing at least one occupied second cache space in response to determining that the second code snippet is executed and determining that at least one second cache space is occupied, to obtain a released second cache space. A loading submodule, for loading a target variable of the second code snippet from the global storage unit to the released second cache space, to obtain at least one data of the second target variable of the second code snippet from the released second cache space.

[0064] In some embodiments, the second execution module includes: a second replacement submodule, configured to replace at least one data of a second target variable of the second code segment in the global storage unit using at least one intermediate result of the loop code block.

[0065] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0066] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.

[0067] Figure 4 A schematic block diagram of an example electronic device 400 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0068] like Figure 4 As shown, the device 400 includes a computing unit 401, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 402 or a computer program loaded from a storage unit 408 to a random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the device 400 can also be stored. The computing unit 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.

[0069] A number of components in the device 400 are connected to the I / O interface 405, including: an input unit 406, such as a keyboard, a mouse, etc.; an output unit 407, such as various types of displays, speakers, etc.; a storage unit 408, such as a disk, an optical disk, etc.; and a communication unit 409, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 409 allows the device 400 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0070] The computing unit 401 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSP), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 401 performs the various methods and processes described above, such as a processing method for loop code. For example, in some embodiments, the processing method for loop code may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 408. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 400 via the ROM 402 and / or the communication unit 409. When the computer program is loaded into the RAM 403 and executed by the computing unit 401, one or more steps of the processing method for loop code described above may be performed. Alternatively, in other embodiments, the computing unit 401 may be configured to execute the processing method for looping code in any other appropriate manner (eg, by means of firmware).

[0071] Various embodiments of the systems and techniques described above herein may be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard parts (ASSPs), system on chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0072] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0073] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories, read-only memories, erasable programmable read-only memories (EPROM) or flash memories, optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0074] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) display or a liquid crystal display (LCD)) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0075] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a Local Area Network (LAN), a Wide Area Network (WAN), and the Internet.

[0076] A computer system may include clients and servers. Clients and servers are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship to each other.

[0077] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.

[0078] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A method for processing loop code, comprising: In response to determining that a total amount of data of a plurality of variables of a loop code block is greater than an available storage space capacity of a local storage unit, storing a plurality of target variables among the plurality of variables in a global storage unit, wherein the loop code block includes a plurality of code fragments, and the local storage unit includes at least one first cache space; In response to determining that a first code snippet among the multiple code snippets is executed, a first target variable of the first code snippet among the multiple target variables is stored in the global storage unit, and determining that there is an idle first cache space in the local storage unit, the first target variable of the first code snippet is stored in the idle first cache space, so that the first code snippet is executed based on the first target variable in the first cache space during the execution of the loop code block.

2. The method according to claim 1, further comprising: In response to determining that a second code snippet among the multiple code snippets is executed, a second target variable of the second code snippet among the multiple target variables is stored in the global storage unit, and determining that there is no free first cache space in the local storage unit, the second code snippet is executed based on the second target variable in the global storage unit during the execution of the loop code block.

3. The method according to claim 1, wherein: The storing of the plurality of target variables into a global storage unit comprises: The plurality of target variables are stored in a plurality of continuous storage spaces of the global storage unit.

4. The method according to claim 1, wherein: The multiple target variables include at least one of the following types of variables: array variables and structure variables.

5. The method according to claim 1, wherein: Executing the first code segment based on the first target variable in the first cache space during the execution of the loop code block includes at least one of the following operations: Acquire at least one data of a first target variable of the first code snippet from the first cache space; At least one data of a first target variable of a first code segment in the first cache space is replaced by using at least one intermediate result of the loop code block.

6. The method according to claim 1, further comprising: In response to determining that the loop code block is executed to completion, writing the processed target variable of the first code snippet to the global storage unit.

7. The method according to claim 2, wherein: The local storage unit also includes at least one second cache space, Executing the second code segment based on the second target variable in the global storage unit during the execution of the loop code block includes: In response to determining that the second code snippet is executed and determining that at least one of the second cache spaces is occupied, releasing at least one of the occupied second cache spaces to obtain a released second cache space; The target variable of the second code snippet is loaded from the global storage unit to the released second cache space, so as to obtain at least one data of the second target variable of the second code snippet from the released second cache space.

8. The method according to claim 2, wherein: Executing the second code segment based on the second target variable in the global storage unit during the execution of the loop code block includes: At least one data of a second target variable of the second code segment in the global storage unit is replaced by at least one intermediate result of the loop code block.

9. A processing device for loop code, comprising: A first storage module, configured to store a plurality of target variables among a plurality of variables in a loop code block to a global storage unit in response to determining that a total amount of data of the plurality of variables in the loop code block is greater than an available storage space capacity of a local storage unit, wherein the loop code block includes a plurality of code fragments, and the local storage unit includes at least one first cache space; A first execution module is used to store the first target variable of the first code snippet to the free first cache space in response to determining that a first code snippet among the multiple code snippets is executed, a first target variable of the first code snippet among the multiple target variables is stored in the global storage unit, and determining that there is a free first cache space in the local storage unit, so as to execute the first code snippet based on the first target variable in the first cache space during the execution of the loop code block.

10. The apparatus according to claim 9, further comprising: A second execution module is used to execute the second code snippet based on the second target variable in the global storage unit during the execution of the loop code block in response to determining that a second code snippet among the multiple code snippets is executed, a second target variable of the second code snippet among the multiple target variables is stored in the global storage unit, and determining that there is no free first cache space in the local storage unit.

11. The device according to claim 9, wherein: The first processing module comprises: The first processing submodule is used to store the plurality of target variables into a plurality of continuous storage spaces of the global storage unit.

12. The device according to claim 9, wherein: The multiple target variables include at least one of the following types of variables: array variables and structure variables.

13. The device according to claim 9, wherein: The first execution module includes at least one of the following operations: an acquisition submodule, configured to acquire at least one data of a first target variable of the first code snippet from the first cache space; The first replacement submodule is used to replace at least one data of a first target variable of a first code segment in the first cache space by using at least one intermediate result of the loop code block.

14. The apparatus according to claim 9, further comprising: A writing module is used for writing the processed target variable of the first code segment into the global storage unit in response to determining that the execution of the loop code block is completed.

15. The device according to claim 10, wherein: The local storage unit also includes at least one second cache space, The second execution module includes: a release submodule, configured to release at least one occupied second cache space in response to determining that the second code fragment is executed and determining that at least one second cache space is occupied, so as to obtain a released second cache space; The loading submodule is used to load the target variable of the second code snippet from the global storage unit to the released second cache space, so as to obtain at least one data of the second target variable of the second code snippet from the released second cache space.

16. The device according to claim 10, wherein: The second execution module includes: The second replacement submodule is used to replace at least one data of the second target variable of the second code fragment in the global storage unit by using at least one intermediate result of the loop code block.

17. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 9.

18. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 9.

19. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Code processing method and device

    CN102929581A

  • Data processing device and method, electronic equipment and storage medium

    CN116701290A

  • Data processing method and device, electronic equipment and storage medium

    CN118426967A

  • Data processing method, processor, computing equipment and device

    CN119322584A

  • Information processing device, compile program, and compile method

    JP2022140995A