Caching method, electronic equipment and storage medium
By determining the clearing priority in the cache based on the application of data in the model, and thus determining the target cache line for clearing, the problem of mistaken clearing in the prior art affects the performance of the model is solved, and the model operation efficiency and performance are improved.
Patent Information
- Application Number
- CN202510239333.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-05-13
AI Technical Summary
The existing cache clearing algorithm is prone to accidentally clear the cache lines required to access the model's subsequent operation, affecting the model's running performance.
The target cache line is determined from each cache line for clearing by the clearing priority determined based on the application of the data stored in the model.
This avoids the problem that the cache line where the data that needs to be used during model operation is accidentally cleared, and improves the model's running efficiency and performance.
Smart Images

Figure CN119988258A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a cache method, electronic equipment and storage medium. Background Art
[0002] The operation of neural network models requires a large amount of data reading and storage operations, which poses a severe challenge to the execution efficiency of memory access. To meet this challenge, the industry generally uses cache to increase the memory access speed, thereby improving the operation efficiency of neural network models.
[0003] The cache uses cache lines as the basic storage unit. When the cache is full, a certain number of cache lines in the cache need to be cleared through the Least Recently Used (LRU) algorithm or the Least Frequently Used (LFU) algorithm to make room for new data.
[0004] The current clearing algorithms are mostly based on the historical access of all cache lines in the cache. The cache lines that need to be cleared are determined from all cache lines without considering the actual situation of the model. This may cause cache lines that need to be accessed in subsequent reasoning to be mistakenly cleared, affecting the performance of the model. Summary of the invention
[0005] The present invention provides a cache method, an electronic device and a storage medium, which are used to solve the defect in the related art that a clearing algorithm is prone to mistakenly clear cache lines that are required to be accessed for subsequent operation of a model.
[0006] The present invention provides a caching method, comprising: Determine a target cache line from each cache line based on the clearing priority of each cache line, wherein the clearing priority is determined based on the application of the data stored in the cache line in the model, wherein the application indicates the application of the data in the process of running the model, and the application is used to reflect the importance of the data in the running of the model; The target cache line is flushed.
[0007] According to a cache method provided by the present invention, the method of determining a target cache line from each cache line based on the clearing priority of each cache line includes: A target cache line is determined from the cache lines based on the clearing priority of each cache line and the historical access situation of each cache line, wherein the historical access situation includes at least one of a historical access time and a historical access count.
[0008] According to a cache method provided by the present invention, the target cache line is determined from each cache line based on the clearing priority of each cache line and the historical access status of each cache line, including: Determining a candidate cache line from each of the cache lines based on historical access conditions of each of the cache lines; The target cache line is determined from the candidate cache lines based on the flushing priorities of the candidate cache lines.
[0009] According to a caching method provided by the present invention, the application situation includes at least one of the number of reuses of data in the model and the application timing.
[0010] According to a caching method provided by the present invention, the determination of the clearing priority includes: Converting the structural diagram of the model into a sequence of one-way operators used for running the model; Determining the priority of each data based on the number of times each operator in the unidirectional operator sequence reuses each data; Determining the time priority of each data based on the application timing of sharing data between operators in the unidirectional operator sequence; Based on the number priority and time priority of each data, the clearing priority of the cache line where each data is located is determined.
[0011] According to a caching method provided by the present invention, the time priority of each data is determined based on the application timing of sharing data between operators in the unidirectional operator sequence, including: Taking the execution time of each operator in the one-way operator sequence as a reference and the preset duration as a range, the time priority of the data shared between each operator and other operators is determined to be the priority corresponding to the preset duration.
[0012] According to a cache method provided by the present invention, the target cache line is determined from each cache line based on the clearing priority of each cache line, and the method also includes: The flushing priority is read from the attributes of the data stored in each of the cache lines.
[0013] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, any of the above-mentioned caching methods is implemented.
[0014] The present invention also provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the caching method described above is implemented.
[0015] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned cache methods.
[0016] The cache method, electronic device and storage medium provided by the present invention can determine the target cache line to be cleared from each cache line by determining the clearing priority based on the application of the data stored in the cache line in the model, thereby avoiding the problem of cache lines containing data that need to be used in model operation being mistakenly cleared, thereby improving the model operation efficiency and optimizing the model performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the present invention or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0018] Figure 1 It is a structural diagram of a neural network model in related technology.
[0019] Figure 2 It is a flow chart of the caching method provided by the present invention.
[0020] Figure 3 It is a flowchart of the method for determining the clearing priority provided by the present invention.
[0021] Figure 4 It is a structural schematic diagram of the cache device provided by the present invention.
[0022] Figure 5 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0023] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0024] As a powerful tool, neural network models play a vital role in image recognition, natural language processing, data mining, etc. However, the operation of such models is highly dependent on a large amount of data reading and storage operations, which poses a severe challenge to the execution efficiency of data reading and storage operations.
[0025] To meet this challenge, the industry generally adopts cache technology to speed up data access, reduce memory latency, and thus improve the operating efficiency of neural network models.
[0026] The cache uses cache lines as the basic storage unit. When the cache is full, that is, when no more data can be cached, it is usually necessary to clear a certain number of cache lines in the cache through the Least Recently Used (LRU) algorithm or the Least Frequently Used (LFU) algorithm to make room for new data.
[0027] Among them, the least recently used algorithm believes that the data that has been accessed recently is more likely to be accessed again in the future, so when the cache space is full, the least recently used algorithm will clear the cache line that has been used least recently. The least frequently used algorithm believes that if a piece of data has been used very few times in the recent period of time, it is also unlikely to be used in the future, so when the cache space is full, the least frequently used algorithm will clear the cache line with the lowest access frequency.
[0028] It can be seen that both the least recently used algorithm and the least frequently used algorithm are based on the historical access of each cache line in the cache, and select the cache line to be cleared from all cache lines in the cache. This means that the currently commonly used clearing algorithms are based on the historical access of the cache line, and do not take into account the actual situation of the neural network model operation.
[0029] For example, Figure 1 It is a structural diagram of a neural network model in related technology. Figure 1 The model structure shown includes a maximum pooling layer (Maxpool), three convolution modules and a splicing module. The three convolution modules are convolution module 1, convolution module 2 and convolution module 3. Each convolution module can be a combination of convolution (Conv), batch normalization (BatchNormalization, BN) and SiLU (Sigmoid-Weighted Linear Unit) layers, so the convolution module can also be called a CBS module, where SiLU is an activation function. In the process of reasoning for the above model structure, the execution order of each module can be the maximum pooling layer → convolution module 1 → convolution module 2 → convolution module 3 → splicing module. In this way, after the execution of convolution module 1, its output tensor is stored in the cache, and during the execution of convolution module 2 and convolution module 3, the output tensor of convolution module 1 will not be accessed. The output tensor of convolution module 1 will not be accessed until the splicing module is executed.
[0030] If the cache is managed and optimized according to the commonly used clearing algorithm mentioned above, the output tensor of convolution module 1 is likely to be mistakenly cleared because it has not been accessed during the execution of convolution modules 2 and 3 and is mistakenly believed to be unlikely to be accessed subsequently. This will directly affect the execution of the splicing module and further affect the operating performance of the above model structure.
[0031] To solve this problem, the present invention provides a caching method. Figure 2 is a flow chart of the caching method provided by the present invention, such as Figure 2 As shown, the method includes: Step 210, based on the clearing priority of each cache line, determine the target cache line from each cache line, the clearing priority is determined based on the application of the data stored in the cache line in the model, the application situation indicates the situation in which the data is applied during the model operation, and the application situation is used to reflect the importance of the data in the model operation.
[0032] Specifically, the cache can be used to store data used in model operation, which can be the input data and output data of each operator in the model, or can be recorded as the input tensor and output tensor of each operator. It can be understood that storing as much data used in model operation as possible and for a long time in the cache can facilitate the computing units used in model operation to read the required data from the memory more quickly and quickly, thereby improving the model operation efficiency and optimizing the model performance.
[0033] When the cache is applied to the storage of application data during model operation, cache optimization can be performed on the cache. Specifically, when the cache is full, it is determined which cache lines in the cache are cleared to provide cache space for new data.
[0034] In the embodiment of the present invention, a clearing priority is configured for each cache line in the cache. Here, the clearing priority is used to reflect the priority of the cache line to be cleared when the cache is full. The higher the clearing priority, the earlier the corresponding cache line is cleared, and the lower the clearing priority, the longer the data stored in the corresponding cache line can be retained.
[0035] The clearing priority is associated with the application of the data stored in the cache line in the model. Here, the application of data in the model refers to the application of data in the process of model operation, which may include the number of times the data is applied, the timing of the data application, etc. Through the application of data in the model, it is possible to analyze whether the data is frequently applied in the model operation, whether it is applied in a short time, or applied after a long time, etc., thereby obtaining the importance of the data in the model operation, and then obtaining the clearing priority referred to in the embodiment of the present invention. For example, the more frequently the data is applied in the model operation and the shorter the application time, the higher the importance of the data in the model operation, and the lower the clearing priority of the cache line storing the data; conversely, the less the data is applied in the model operation and the longer it is not applied, the lower the importance of the data in the model operation and the higher the clearing priority of the cache line storing the data.
[0036] When performing cache optimization, the target cache line can be selected based on the clearing priority of each cache line. For example, the cache line with the highest clearing priority among the cache lines can be used as the target cache line to be cleared first; or, the cache line to be cleared determined by the clearing algorithm in the relevant technology can be applied first, and then the cache line with the highest clearing priority can be selected from the cache lines to be cleared as the target cache line; or, the clearing score of each cache line determined by the clearing algorithm in the relevant technology and the clearing priority in the embodiment of the present invention can be weighted, and the target cache line can be determined from each cache line based on the weighted result, which is not specifically limited in the embodiment of the present invention.
[0037] Step 220: clear the target cache line.
[0038] Specifically, after determining the target cache line, the target cache line can be cleared, thereby providing cache space for new data. It can be understood that, since the target cache line to be cleared in the embodiment of the present invention is determined based on the clearing priority, and the clearing priority is associated with the application of the data in the model, the application of the data in the model operation is actually considered during the cache optimization process, which can avoid the problem that relatively important data in the model operation is cleared from the cache, resulting in affecting the model operation performance.
[0039] In the method provided in the embodiment of the present invention, by determining the clearing priority based on the application of the data stored in the cache line in the model, the target cache line is determined from each cache line for clearing, thereby avoiding the problem of cache lines containing data needed in model operation being mistakenly cleared, thereby improving the model operation efficiency and optimizing the model performance.
[0040] Based on the above embodiment, in step 210, determining the target cache line from each cache line based on the clearing priority of each cache line includes: A target cache line is determined from the cache lines based on the clearing priority of each cache line and the historical access situation of each cache line, wherein the historical access situation includes at least one of a historical access time and a historical access count.
[0041] Specifically, for any cache line, the historical access situation of the cache line, that is, the situation in which the cache line was previously accessed, may specifically include at least one of the historical access time and the historical access count. Among them, the historical access time is the time when the cache line was previously accessed, which may specifically be the moment when the cache line was last accessed, or it may be the length of time from the last access of the cache line to the current moment, etc. It is understandable that the longer the historical access time is, the longer the cache line has not been accessed, and the smaller the probability of subsequent access; the historical access count is the number of times the cache line has been accessed before, which may specifically be the cumulative number of times the cache line has been accessed after data is written to the cache line, or it may be the cumulative number of times the cache line has been accessed within a preset period of time, etc. It is understandable that the more historical access times there are, the greater the probability that the cache line will be accessed later.
[0042] Therefore, when determining the target cache line from each cache line, not only the flushing priority of each cache line can be referred to, but also the historical access situation of each cache line can be referred to. In specific implementation, the flushing priority of each cache line and the historical access situation of each cache line can be combined to calculate the probability of each cache line being the target cache line. For example, for any cache line, the flushing priority and the historical access situation of the cache line can be input into a pre-built algorithm to obtain the probability of the cache line being the target cache line. For another example, for any cache line, the score corresponding to the flushing priority of the cache line and the score corresponding to the historical access situation of the cache line can be weighted summed, and the probability of the cache line being the target cache line can be determined based on the sum of the scores; or, the cache line with the highest flushing priority can be selected first, and in the case of multiple cache lines with the highest flushing priority, the historical access situation of the cache line with the highest flushing priority can be further applied to select the target cache line; or, cache line selection can be performed based on the historical access situation of each cache line first, and in the case of multiple cache lines selected, the cache line with the highest flushing priority can be selected as the target cache line. The embodiment of the present invention does not specifically limit this.
[0043] In the method provided in the embodiment of the present invention, the clearing priority of the cache line and the historical access situation of the cache line are combined, and the target cache line is determined based on the application of data in the model and the access situation of the cache, thereby further ensuring the rationality and reliability of cache optimization.
[0044] Based on any of the above embodiments, in step 210, determining the target cache line from each cache line based on the clearing priority of each cache line and the historical access status of each cache line includes: Determining a candidate cache line from each of the cache lines based on historical access conditions of each of the cache lines; The target cache line is determined from the candidate cache lines based on the flushing priorities of the candidate cache lines.
[0045] Specifically, when determining the target cache line from each cache line, cache line screening can be performed based on the historical access of each cache line. In an embodiment of the present invention, cache lines that may be cleared are first screened based on the historical access of each cache line and recorded as candidate cache lines.
[0046] Here, cache line screening is performed based on the historical access situation of each cache line, which can be implemented through clearing algorithms such as the lowest frequency use algorithm and the least recently used algorithm. For example, the least recently used algorithm implemented by the age method is used to select candidate cache lines. Specifically, a field named age can be set in the status bit interval of each cache line. Every time the cache is accessed, the value of the age field of the cache line that has not been accessed is automatically reduced by 1 until the value of the age field is 0. It is understandable that if the value of the age field of a cache line is 0, it means that the cache line has not been accessed for a long time, and the cache line can be cleared first when the cache block is replaced next time. That is, the cache line with the value of 0 in the age field can be used as a candidate cache line. It is understandable that the number of candidate cache lines can be one or more.
[0047] When there is only one candidate cache line, the candidate cache line can be directly used as the target cache line; when there are multiple candidate cache lines, the target cache line can be further selected from the candidate cache lines based on the clearing priority of each candidate cache line.
[0048] For example, based on the least recently used algorithm implemented by the age method to select candidate cache lines, a field can be added to the status bit interval of each cache line, and the newly added field is used to indicate the flushing priority of the cache line. When a cache block replacement occurs, if there are multiple candidate cache lines with an age field value of 0, the candidate cache line with the highest flushing priority field value can be selected from these candidate cache lines as the target cache line for flushing.
[0049] Furthermore, if there are multiple candidate cache lines with the highest value of the clearing priority field, the candidate cache line with the lowest cache address may be selected from the multiple candidate cache lines as the target cache line.
[0050] Based on any of the above embodiments, the application situation includes at least one of the number of times the data is reused in the model and the application timing.
[0051] Specifically, for any cache line, the clearing priority of the cache line is determined based on the application of the data stored in the cache line in the model. The application referred to here may include at least one of the number of reuses and application timing of the data in the model.
[0052] The number of times data is reused in the model refers to the total number of times the data is applied in one model run. The number of reuses can reflect whether the data is frequently used in the model run, and thus determine whether the data is important in the model run.
[0053] The application timing of data in the model refers to the time when the data is applied in a model run. The application timing can reflect whether the data is applied in a short time or a long time in the model run, which can be used to determine whether the data is important in the model run.
[0054] In the method provided in the embodiment of the present invention, the clearing priority of the cache line is determined by at least one of the number of times the data is reused in the model and the application timing to achieve cache optimization. During the cache optimization process, the application of the data during model operation can be actually considered, thereby ensuring the model operation performance.
[0055] Based on any of the above embodiments, Figure 3 is a flow chart of a method for determining clearing priority provided by the present invention, such as Figure 3 As shown, the determination of the clearing priority includes: Step 310: convert the structural diagram of the model into a one-way operator sequence used for model operation.
[0056] Specifically, the structure diagram of a model is an intuitive representation of the model, and the structure diagram can describe the relationship between various components of the model (such as layers, connections, etc.) and the direction of data flow.
[0057] After obtaining the structure graph of the model, the structure graph of the model can be converted into a one-way operator sequence. The conversion from the structure graph to the one-way operator sequence can be understood as linearizing the structured computational graph that may contain branches and loops into a list of operators that are executed sequentially. For example, Figure 1The structural diagram shown is converted into a unidirectional operator sequence in the form of “maximum pooling layer→convolution module 1→convolution module 2→convolution module 3→splicing module”.
[0058] Here, the structural graph is converted into a one-way operator sequence, which can be implemented through a deep learning framework such as PyTorch.
[0059] Step 320: Determine the priority of each data based on the number of times each operator in the unidirectional operator sequence reuses each data.
[0060] Specifically, after obtaining the one-way operator sequence, the number of data reuses can be counted for each operator in the one-way operator sequence. Here, for any operator in the one-way operator sequence, the number of data reuses for the operator is counted, specifically, the number of times each data is used is recorded during the operation of the operator. What can be obtained is the number of times each data is reused during the operation of each operator, and it is used as the number of times each data is reused in the model. If there is a situation where a data is used by multiple operators, the number of times the data is reused during the operation of multiple operators can be accumulated to obtain the number of times the data is reused in the model.
[0061] After obtaining the number of reuses of each data in the model, the priority of each data can be determined based on the number of reuses. The priority of the number of reuses here is used to reflect the priority of clearing the cache line where the data is located based on the number of reuses of the data.
[0062] Specifically, the priority of the number of times can be determined according to the number of times of multiplexing. For example, the lowest priority of the number of times can be set for the data with the highest number of times of multiplexing, and the highest priority of the number of times can be set for the data with the lowest number of times of multiplexing. For another example, the lowest priority of the number of times can be set for the data with the highest number of times of multiplexing, and the highest priority of the number of times can be set for the remaining data.
[0063] Step 330: Determine the time priority of each data based on the application timing of sharing data between operators in the unidirectional operator sequence.
[0064] Specifically, after obtaining the one-way operator sequence, the application timing of the shared data can be analyzed according to whether there is shared data between the operators in the one-way operator sequence. Here, for two operators, whether there is shared data between the two operators can be understood as whether the data applied by the two operators are the same data, that is, whether the data applied by one operator is also applied in the other operator.
[0065] In the case of shared data, the application timing of the shared data can be obtained. Here, the application timing of shared data refers to the time interval between two operators sharing data, for example, Figure 1 In the figure, the output tensor of convolution module 1 is used as the input tensor of the splicing module. In the unidirectional operator sequence "max pooling layer → convolution module 1 → convolution module 2 → convolution module 3 → splicing module", the execution time of convolution module 1 is T, and the execution time of the splicing module is T+3, so the application time of the output tensor of convolution module 1 is T+3.
[0066] After obtaining the application timing of each data, the time priority of each data can be determined. It is understandable that for data that is not shared between operators during model operation, its application timing can be set to null. The time priority thus obtained is used to reflect the priority of the cache line where the data is located to be cleared based on the application timing of the data.
[0067] Specifically, the time priority of each data can be determined according to whether each data is shared in the model and the application timing of the shared data. For example, the lowest time priority can be set for data with an application timing in the range of T-3 to T+3, and the highest time priority can be set for other data. For another example, the lowest time priority can be set for data with an application timing closest to T, and the highest time priority can be set for data with an empty application timing.
[0068] It should be noted that the embodiment of the present invention does not limit the execution order of step 320 and step 330 , and step 320 may be executed before or after step 330 , or may be executed synchronously with step 330 .
[0069] Step 340: Determine the clearing priority of the cache line where each data is located based on the number priority and time priority of each data.
[0070] Specifically, after respectively obtaining the number priority and the time priority of each data, the number priority and the time priority of each data may be combined as the clearing priority of the cache line where each data is located.
[0071] For example, for any data, the number priority of the data is P0 and the time priority is P2, then the clearing priority of the cache line where the data is located can be recorded as P02.
[0072] Based on any of the above embodiments, in step 330, determining the time priority of each data based on the application timing of sharing data between operators in the unidirectional operator sequence includes: Taking the execution time of each operator in the one-way operator sequence as a reference and the preset duration as a range, the time priority of the data shared between each operator and other operators is determined to be the priority corresponding to the preset duration.
[0073] Specifically, the preset duration is a pre-set duration. There may be multiple preset durations. Different combinations of preset durations and operator execution times may form different time intervals. For example, the preset duration may be 1, 2, or 3. Assuming that the execution time of an operator is T, then taking T as the reference and the preset duration as the range, the time intervals that can be obtained include [T-1, T +1], [T-2, T +2], and [T-3, T +3].
[0074] After determining the time interval, the data shared in the time zone can be determined, and then the time priority of the data can be determined. The time priority determined here is related to the time interval in which the data is shared, and the smaller the time interval, the closer the data sharing occurs, and the lower the time priority, and the larger the time interval, the farther the data sharing occurs, and the higher the time priority.
[0075] For example, the time priority of data shared within [T-1, T +1] is lower than the time priority of data shared within [T-3, T +3].
[0076] The method provided in the embodiment of the present invention determines the time priority of data according to the priority corresponding to the preset duration, thereby realizing the priority subdivision at the application timing level, and helping to improve the reliability of cache optimization.
[0077] Based on any of the above embodiments, in step 210, the target cache line is determined from each cache line based on the clearing priority of each cache line, and the method also includes: The flushing priority is read from the attributes of the data stored in each of the cache lines.
[0078] Specifically, for each cache line, the clearing priority of the cache line can be obtained from the attributes of the stored data, that is, the clearing priority of the cache line is actually the clearing priority of the stored data, and the clearing priority of the data is a type of data attribute.
[0079] The clearing priority of data can be passed in as a type of data attribute when creating the data; or, the data ID and the data clearing priority can be included in a tuple, and the tuple is received through the operator interface to obtain the clearing priority therein as a type of attribute of the data corresponding to the ID. The embodiment of the present invention does not make any specific limitation on this.
[0080] After the flushing priority is read from the attribute of the data, the flushing priority can be transmitted to the cache controller to facilitate cache optimization.
[0081] Based on any of the above embodiments, the caching method may include: Assume that in a model operation, there are three data A, B, and C, and based on the network structure of the model, the application of the above three data in the model is determined as follows: data A will be read twice in the near future, that is, the number of reuses of data A is 2, and the application time is relatively recent; data B will be read repeatedly in the current operator in the model, that is, the number of reuses of data B>2, and the application time is recent; data C is the output of the current operator in the model, and will not be accessed repeatedly for a long time, that is, the number of reuses of data C is 0, and the application time is relatively far. Based on the number of reuses and application time of each data, the data can be arranged from high to low according to their importance in the model operation, that is, data B, data A, data C, and accordingly, the priority of each data can be determined, with data A having a medium priority, data B having a low priority, and data C having a high priority.
[0082] Then, you can configure the clearing priority of each data in the properties of each data.
[0083] During the model operation, the clearing priority of the data can be read from the attributes of the data, and when the data is written into the cache line, the clearing priority of the data is also written into the status field of the cache line. When managing the cache, the clearing priority in the cache line status field and the clearing algorithms such as the lowest frequency use algorithm and the least recently used algorithm can be combined to determine the target cache line from each cache line for clearing. For example, the cache line with the highest clearing priority can be selected first, and when multiple cache lines with the highest priority are selected, the scores calculated for each cache line by the clearing algorithms such as the lowest frequency use algorithm and the least recently used algorithm are compared, thereby further selecting the target cache line from the multiple cache lines with the highest priority; for another example, the scores calculated for each cache line by the clearing algorithms such as the lowest frequency use algorithm and the least recently used algorithm can be multiplied by the coefficient corresponding to the clearing priority as the score for treating each cache line as the target cache line, thereby selecting the target cache line. Here, the higher the clearing priority, the higher the corresponding coefficient, and the higher the possibility that the cache line is regarded as the target cache line for clearing.
[0084] The cache device provided by the present invention is described below. The cache device described below and the cache method described above can be referred to each other.
[0085] Figure 4 is a schematic diagram of the structure of the cache device provided by the present invention, such as Figure 4 As shown, the device comprises: A selection unit 410 is used to determine a target cache line from each cache line based on a clearing priority of each cache line, wherein the clearing priority is determined based on an application of data stored in the cache line in the model, wherein the application indicates a situation in which the data is applied during the operation of the model, and the application is used to reflect the importance of the data in the operation of the model; The flushing unit 420 is configured to flush the target cache line.
[0086] In the device provided in the embodiment of the present invention, by determining the clearing priority based on the application of the data stored in the cache line in the model, the target cache line is determined from each cache line for clearing, thereby avoiding the problem of cache lines containing data that need to be used in model operation being mistakenly cleared, thereby improving the model operation efficiency and optimizing the model performance.
[0087] Based on any of the above embodiments, the selection unit is specifically used for: A target cache line is determined from the cache lines based on the clearing priority of each cache line and the historical access situation of each cache line, wherein the historical access situation includes at least one of a historical access time and a historical access count.
[0088] Based on any of the above embodiments, the selection unit is specifically used for: Determining a candidate cache line from each of the cache lines based on historical access conditions of each of the cache lines; The target cache line is determined from the candidate cache lines based on the flushing priorities of the candidate cache lines.
[0089] Based on any of the above embodiments, the application situation includes at least one of the number of times the data is reused in the model and the application timing.
[0090] Based on any of the above embodiments, the device further includes a priority unit, configured to: Converting the structural diagram of the model into a sequence of one-way operators used for model operation; Determining the priority of each data based on the number of times each operator in the unidirectional operator sequence reuses each data; Determining the time priority of each data based on the application timing of sharing data between operators in the unidirectional operator sequence; Based on the number priority and time priority of each data, the clearing priority of the cache line where each data is located is determined.
[0091] Based on any of the above embodiments, the priority unit is specifically used for: Taking the execution time of each operator in the one-way operator sequence as a reference and the preset duration as a range, the time priority of the data shared between each operator and other operators is determined to be the priority corresponding to the preset duration.
[0092] Based on any of the above embodiments, the priority unit is further configured to: The flushing priority is read from the attributes of the data stored in each of the cache lines.
[0093] Figure 5 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 5 As shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530 and a communication bus 540, wherein the processor 510, the communication interface 520 and the memory 530 communicate with each other through the communication bus 540. The processor 510 may call the logic instructions in the memory 530 to execute the cache method, which includes: Determine a target cache line from each cache line based on the clearing priority of each cache line, wherein the clearing priority is determined based on the application of the data stored in the cache line in the model, wherein the application indicates the application of the data in the process of running the model, and the application is used to reflect the importance of the data in the running of the model; The target cache line is flushed.
[0094] In addition, the logic instructions in the above-mentioned memory 530 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, or the part that contributes to the relevant technology or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0095] On the other hand, the present invention further provides a computer program product, the computer program product includes a computer program, the computer program can be stored in a non-transitory computer-readable storage medium, when the computer program is executed by a processor, the computer can execute the cache method provided by the above methods, the method includes: Determine a target cache line from each cache line based on the clearing priority of each cache line, wherein the clearing priority is determined based on the application of the data stored in the cache line in the model, wherein the application indicates the application of the data in the process of running the model, and the application is used to reflect the importance of the data in the running of the model; The target cache line is flushed.
[0096] In another aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the caching method provided by the above methods is implemented, and the method includes: Determine a target cache line from each cache line based on the clearing priority of each cache line, wherein the clearing priority is determined based on the application of the data stored in the cache line in the model, wherein the application indicates the application of the data in the process of running the model, and the application is used to reflect the importance of the data in the running of the model; The target cache line is flushed.
[0097] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0098] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the relevant technology can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiment.
[0099] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A caching method, characterized in that: include: Determine a target cache line from each cache line based on the clearing priority of each cache line, wherein the clearing priority is determined based on the application of the data stored in the cache line in the model, wherein the application indicates the application of the data in the process of running the model, and the application is used to reflect the importance of the data in the running of the model; The target cache line is flushed.
2. The caching method according to claim 1, characterized in that: The step of determining a target cache line from each cache line based on the clearing priority of each cache line includes: A target cache line is determined from the cache lines based on the clearing priority of each cache line and the historical access situation of each cache line, wherein the historical access situation includes at least one of a historical access time and a historical access count.
3. The caching method according to claim 2, characterized in that: The determining the target cache line from each cache line based on the clearing priority of each cache line and the historical access status of each cache line includes: Determining a candidate cache line from each of the cache lines based on historical access conditions of each of the cache lines; The target cache line is determined from the candidate cache lines based on the flushing priorities of the candidate cache lines.
4. The caching method according to claim 1, characterized in that: The application situation includes at least one of the number of times the data is reused in the model and the application timing.
5. The caching method according to claim 4, characterized in that: The determination of the clearing priority includes: Converting the structural diagram of the model into a sequence of one-way operators used for running the model; Determining the priority of each data based on the number of times each operator in the unidirectional operator sequence reuses each data; Determining the time priority of each data based on the application timing of sharing data between operators in the unidirectional operator sequence; Based on the number priority and time priority of each data, the clearing priority of the cache line where each data is located is determined.
6. The caching method according to claim 5, characterized in that: The determining the time priority of each data based on the application timing of the data shared between the operators in the unidirectional operator sequence includes: Taking the execution time of each operator in the one-way operator sequence as a reference and the preset duration as a range, the time priority of the data shared between each operator and other operators is determined to be the priority corresponding to the preset duration.
7. The caching method according to any one of claims 1 to 6, characterized in that: The step of determining a target cache line from each cache line based on the clearing priority of each cache line also includes: The flushing priority is read from the attributes of the data stored in each of the cache lines.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the cache method according to any one of claims 1 to 7 is implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the cache method according to any one of claims 1 to 7 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the cache method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Cache line data multiplexing method, electronic equipment and storage medium
CN121051039A