System-on-chip caching method and system adopting adaptive prediction technology
By using adaptive prediction technology and label instructions management buffers in system-level chip cache, the long waiting time and resource waste problems during jump instruction processing are solved, and the effect of improving the overall efficiency of the system and resource utilization is achieved.
Patent Information
- Application Number
- CN202510244052.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-03
AI Technical Summary
When processing jump instructions, the prior art needs to clear the cache and re-acquire the instructions from the main memory, resulting in a long wait time, reducing the overall efficiency of the system, and increasing resource waste.
Adaptive prediction technology is adopted to manage instruction flow in the buffer through tag instructions, predict jump targets and cache them in advance, avoiding clearing caches and waiting for main memory to reload instructions.
It significantly reduces the waiting time when executing jump instructions, improves the overall system efficiency, reduces resource redundancy, and improves resource utilization.
Smart Images

Figure CN120144183A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of electronic information technology, and particularly relates to a system-level chip caching method and system adopting an adaptive prediction technology. Background Art
[0002] With the continuous progress of integrated circuits and industrial intelligence, the design scale of embedded SOCs is expanding day by day, and the requirements for performance are also increasing day by day. This makes the comprehensive performance and execution efficiency of instruction processing face higher challenges. Especially for DSPs, in order to improve performance, the currently commonly adopted solution is to add a level of cache after Flash to make up for the speed difference between the high-speed CPU and the low-speed Flash, thereby improving the comprehensive efficiency of system instruction execution to a certain extent. However, this technology has obvious deficiencies in processing jump instructions: it will clear the cache and re-fetch instructions from the Flash address after the jump, which will generate a long waiting time and thus seriously reduce the comprehensive efficiency of the system. Since the proportion of jump instructions in the program is relatively high, this problem is particularly prominent. To further solve this problem, a first buffer and a second buffer are added between the main memory and the decoding unit in the prior art. By using the second buffer to pre-cache the to-be-executed instructions after the jump, the system does not need to wait for the instructions to be written into the buffer, thereby effectively improving the comprehensive efficiency. However, since the cost of the buffer is generally high, this solution actually only uses one of the two buffers to make up for the speed difference between the high-speed CPU and the low-speed Flash, and the other is in a standby state most of the time, resulting in a certain waste of resources. Summary of the Invention
[0003] The present invention aims to solve at least one of the technical problems existing in the prior art. For this purpose, the present invention proposes a system-level chip caching method and system adopting an adaptive prediction technology. The aim is to improve the resource utilization rate while improving the comprehensive efficiency.
[0004] In a first aspect, an embodiment of the present invention provides a system-level chip caching method adopting an adaptive prediction technology. The method includes the following steps: Read a plurality of consecutive first to-be-executed instructions from the main memory according to a preset rule and cache them in a buffer. The buffer also includes a first label indication for indicating the number of remaining unexecuted instructions in the first to-be-executed instructions; Based on the first label indication, read a first target to-be-executed instruction in the first to-be-executed instructions from the buffer to a decoding unit, and the decoding unit determines the instruction type of the first target to-be-executed instruction. The instruction type includes a jump type and a non-jump type; When it is determined that the instruction type of the first target instruction to be executed is a jump type, a jump prediction message is sent to the adaptive prediction module; the adaptive prediction module makes a prediction based on the jump prediction message and sends the prediction result to the address selection module; The address selection module reads a plurality of consecutive second instructions to be executed from the main memory and caches them in the buffer based on the prediction result. The second instructions to be executed are the instructions to be executed after the jump. The buffer also includes a second label indication for indicating the number of remaining unexecuted instructions in the second instructions to be executed; When the decoding unit determines to execute the first target instruction to be executed, a jump execution message is sent to the buffer; after receiving the jump execution message, the buffer reads the second target instruction to be executed in the second instructions to be executed to the decoder based on the second label indication.
[0005] In a second aspect, a system-level chip cache system adopting an adaptive prediction technology is provided, including: a main memory, a buffer, a decoding unit, an adaptive prediction module, and an address selection module; The main memory is configured to read a plurality of consecutive first instructions to be executed according to a preset rule and cache them in the buffer. The buffer also includes a first label indication for indicating the number of remaining unexecuted instructions in the first instructions to be executed; The buffer is configured to read the first target instruction to be executed in the first instructions to be executed to the decoding unit based on the first label indication; The decoding unit is configured to determine the instruction type of the first target instruction to be executed. The instruction type includes a jump type and a non-jump type; when it is determined that the instruction type of the first target instruction to be executed is a jump type, a jump prediction message is sent to the adaptive prediction module; The adaptive prediction module is configured to make a prediction based on the jump prediction message and send the prediction result to the address selection module; The address selection module is configured to instruct the main memory to read a plurality of consecutive second instructions to be executed and cache them in the buffer based on the prediction result. The second instructions to be executed are the instructions to be executed after the jump. The buffer also includes a second label indication for indicating the number of remaining unexecuted instructions in the second instructions to be executed; The decoding unit is further configured to send a jump execution message to the buffer when it is determined to execute the first target instruction to be executed; The buffer is further configured to read the second target instruction to be executed in the second instructions to be executed to the decoder based on the second label indication after receiving the jump execution message.
[0006] The system - level chip caching method and system adopting an adaptive prediction technology according to an embodiment of the present invention have at least the following technical effects: It is not necessary to clear the instructions to be executed that have been cached in the buffer, nor is it necessary to have more buffers. Jumps are achieved through tag indication, and it is not necessary to let the CPU wait for the main memory to rewrite the instructions to be executed after the jump into the buffer. Compared with traditional caching methods, the waiting time when executing jump instructions is greatly reduced, the overall system efficiency is effectively improved, resource redundancy is reduced, and resource utilization is enhanced. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0008] Figure 1 FIG. is a schematic structural diagram of a system - level chip caching system adopting an adaptive prediction technology provided by an embodiment of the present application; Figure 2 FIG. is a schematic structural diagram of another system - level chip caching system adopting an adaptive prediction technology provided by an embodiment of the present application; Figure 3 FIG. is a schematic flowchart of a system - level chip caching method adopting an adaptive prediction technology provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0009] The following will clearly and completely describe the technical solutions of the present invention in conjunction with the embodiments. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0010] To better describe the system - level chip caching method and system adopting an adaptive prediction technology according to an embodiment of the present invention, an embodiment of the present application provides an architecture platform for executing the system - level chip caching method adopting an adaptive prediction technology, as Figure 1 shown in FIG. 2, the system - level chip caching system adopting an adaptive prediction technology includes: a main memory, a buffer, a decoding unit, an adaptive prediction module, and an address selection module; Among them, the main memory (Main Memory): refers to the component in a computer for storing data and programs, which can be directly accessed by the CPU and is used to temporarily store the running programs and data.
[0011] Buffer: A storage device used to temporarily store data to coordinate data transfer between devices with different speeds or processing capabilities.
[0012] Decoding Unit: Usually refers to the instruction decoder, which is responsible for parsing the instructions received by the CPU and converting them into control signals that the computer can understand.
[0013] Adaptive Prediction Module: Used to predict system behavior or results, and self-adjust according to historical data and patterns to improve the accuracy of prediction.
[0014] Address Selection Module: Responsible for selecting an address from multiple possible addresses for reading or writing data.
[0015] In the embodiment of the present application, the main memory is used to read multiple consecutive first to-be-executed instructions according to a preset rule and cache them in the buffer.
[0016] The preset rule is a preset instruction reading order or logical condition, which is used to determine the call priority and address sequence of the to-be-executed instructions in the main memory.
[0017] Consecutive instructions refer to the physical or logical adjacent storage state of the first to-be-executed instructions in the main memory, ensuring that the instruction stream is read without interruption.
[0018] The buffer is used to store multiple consecutive first to-be-executed instructions. The buffer also includes a first label indication. The first label indication is used to indicate the number of remaining unexecuted instructions in the first to-be-executed instructions. The buffer is used to read the first target to-be-executed instruction in the first to-be-executed instructions to the decoding unit based on the first label indication; The decoding unit is used to determine the instruction type of the first target to-be-executed instruction. The instruction type includes jump type and non-jump type; when it is determined that the instruction type of the first target to-be-executed instruction is jump type, the decoding unit sends a jump prediction message to the adaptive prediction module.
[0019] Among them, jump type instructions determine whether to change the program counter (PC) through the operation code to implement branches, loops, or function calls; non-jump type instruction operation codes do not modify the PC and sequentially execute data processing, storage, and other operations.
[0020] The adaptive prediction module is used to make a prediction based on the jump prediction message and send the prediction result to the address selection module.
[0021] The address selection module is configured to instruct the main memory to read multiple consecutive second to-be-executed instructions into the buffer based on the prediction result, where the second to-be-executed instructions are the to-be-executed instructions after the jump.
[0022] The address selection template can parse the original address text and the prediction message packet, extract the instruction storage address features (segment address / offset / cache line identifier); generate a candidate address set according to the jump prediction result, and establish an instruction prefetch queue: main memory → L2 cache → L1 instruction cache; verify by grading according to the memory physical structure: NUMA node → memory channel → Bank group → row address; cache line alignment optimization (64-byte boundary detection) can be implemented. Conflict handling can also be performed, for example, initiating a page table traversal or triggering a page fault; executing a cache replacement policy (LRU / Pseudo-LRU).
[0023] The buffer further includes a second tag indication for indicating the number of remaining unexecuted instructions in the second to-be-executed instructions; The decoding unit is further configured to send a jump execution message to the buffer when it is determined to execute the first target to-be-executed instruction; The buffer is further configured to read the second target to-be-executed instruction in the second to-be-executed instructions into the decoder based on the second tag indication after receiving the jump execution message.
[0024] Through the embodiments of the present application, there is no need to clear the to-be-executed instructions already cached in the buffer, nor is there a need for more buffers. Jumping is achieved through tag indication, and there is no need to let the CPU wait for the main memory to rewrite the to-be-executed instructions after the jump into the buffer. Compared with the traditional caching method, the waiting time when executing jump instructions is greatly reduced, the overall system efficiency is effectively improved, resource redundancy is reduced, and resource utilization is improved.
[0025] In an embodiment of the present application, the main memory is responsible for reading a series of consecutive instructions and storing them in a buffer. The buffer is used to temporarily store these instructions and contains a label indicator that shows the number of unexecuted instructions. When the decoding unit receives these instructions, it identifies the type of the instructions and determines whether it is a jump instruction. If it is a jump instruction, the decoding unit sends the jump prediction information to the adaptive prediction module, which makes a jump decision based on the prediction information and passes the result to the address selection module. The address selection module guides the main memory to read the instruction sequence after the jump according to the prediction result and stores it in the buffer. The buffer also contains another label indicator for showing the number of unexecuted instructions in the instruction sequence after the jump. When the decoding unit determines to execute the jump instruction, it sends jump execution information to the buffer, and the buffer then reads the target instruction after the jump to the decoding unit according to this information. This embodiment manages the jump of instructions by using label indicators, avoiding the need to empty the buffer or wait for the main memory to reload instructions when executing jump instructions, thus significantly reducing the waiting time, improving the system efficiency, reducing resource waste, and increasing the utilization rate of resources.
[0026] In an embodiment of the present application, the label indicator may include various implementation manners. As an example, as Figure 2 shown, the buffer includes a plurality of labels, each label corresponding to a storage area, and each area is used to store the first instruction to be executed or the second instruction to be executed read at the same time; Each label respectively includes: a quantity identifier, an activation status identifier, and a jump identifier.
[0027] The label with the first state of the jump identifier corresponds to the first instruction to be executed, and the label with the second state of the jump identifier corresponds to the second instruction to be executed.
[0028] The quantity identifier is used to record the number of unexecuted instructions in the corresponding area. When the number of unexecuted instructions corresponding to the label is 0, the label is used to record that the corresponding area can be covered.
[0029] The activation status identifier is used to indicate whether the instruction to be executed stored in the corresponding area can be read.
[0030] Among them, the first label indicator corresponds to the label with the first state of the jump identifier being in the active state and the label with the second state of the jump identifier being in the non-active state.
[0031] The second label indicator corresponds to the label with the first state of the jump identifier being in the non-active state and the label with the second state of the jump identifier being in the active state.
[0032] For example, when the decoding unit recognizes a jump instruction, the jump identifier in the buffer switches from the first state (corresponding to the first instruction to be executed) to the second state (corresponding to the second instruction to be executed), and the quantity identifier of the second instruction to be executed is greater than zero; at the same time, the activation status identifier of the original first instruction to be executed changes from valid to invalid, but its quantity identifier is retained for subsequent overwrite determination.
[0033] After the main memory loads the second instruction sequence to be executed according to the jump prediction result of the address selection module, the buffer marks the switching of the old and new instruction streams through the change of the jump identifier status, and the quantity identifier continuously tracks the number of unexecuted instructions. When the decoding unit confirms the execution of the jump, the activation status identifier is updated synchronously, enabling the second instruction to be executed to be read, while retaining the unexhausted part of the first instruction stream for exception rollback.
[0034] As another example, as Figure 1 shown, the buffer may include multiple tags, each tag for recording an instruction to be executed; the first tag is used to record the first unexecuted instruction among the first instructions to be executed; the second tag is used to record the first unexecuted instruction among the second instructions to be executed; the third tag is used to record whether it is the first instruction to be executed or the second instruction to be executed that is currently being executed; wherein, the first tag indicates that the third tag indicates the execution of the first instruction to be executed; the second tag indicates that the third tag indicates the execution of the second instruction to be executed.
[0035] After determining to execute the first instruction to be executed or the second instruction to be executed according to the third tag, the specific instruction to be executed is determined according to the first tag or the second tag.
[0036] For example, the quantity identifier of the unexecuted instructions in the storage area corresponding to a certain tag in the buffer is zeroed, and its activation status identifier is invalid (not the current instruction stream being executed); at the same time, the buffer detects that the main memory needs to load a new instruction sequence and there is no free storage area. When the third tag indicates switching to the second instruction stream to be executed, the tags corresponding to the original first instruction stream continuously monitor the quantity identifier. If all the instructions in this area have been executed and not reactivated, it is marked as a coverable state. When the buffer is full and needs to write new instructions, the area corresponding to the coverable tag is preferentially selected for replacement, and at the same time, the jump identifier and the quantity identifier are updated to ensure that the continuity of the instructions is not interrupted.
[0037] In addition, this embodiment also provides a dynamic management mechanism for tag indication. During the execution of instructions, the tag indication is dynamically updated according to the execution status of the instructions. For example, when the instructions in an area are executed, the quantity identifier is updated accordingly, indicating that the number of unexecuted instructions in this area has decreased.
[0038] Meanwhile, the activation status identifier also changes according to the execution of the instructions. Once the instructions are executed, the corresponding storage area will be marked as overwriteable. At this time, the quantity identifier will display 0, and the activation status identifier will be updated to the inactive state, indicating that the instructions in this area no longer need to be retained.
[0039] During the execution of the instructions, if a jump is required, the jump identifier will play a key role. It will, according to the logical requirements of the program, direct the processor to jump to the storage area pointed to by the corresponding label to obtain the next instruction to be executed.
[0040] To ensure the correctness and efficiency of the instruction execution, this embodiment also introduces a verification mechanism for label indication. Before each instruction jump or execution, the system will check the label indication to ensure that it correctly reflects the status of the storage area.
[0041] In addition, this embodiment also considers exception handling. During the execution of the instructions, if an exception occurs, the system will quickly locate the problem according to the label indication and take corresponding recovery measures to ensure the stable operation of the system.
[0042] In some embodiments, the decoding unit includes a pre-decoder and a multi-stage decoder connected in sequence; Among them, the pre-decoder is used to confirm the instruction type of the first target instruction to be executed read; if the first target instruction to be executed is confirmed as a jump instruction type, generate a jump prediction message; the multi-stage decoder is used to confirm the first jump requirement status of the first target instruction to be executed, and the first jump requirement status includes a first jump state and a first non-jump state; if the first jump requirement status is confirmed as the first jump state, generate a jump execution message.
[0043] As an example, this solution adopts a dual verification mechanism of a first-level instruction decoder and a second-level privilege decoder. The first-level decoder first parses the instruction format, verifies the validity of the operation code and whether the target address is within the range allowed by the instruction set architecture; the second-level decoder detects the access privilege of the target address and compares the current privilege level with the access privilege mark of the memory area where the target address is located. Only when both levels of decoders output a verification passed signal, will the control unit generate a jump execution message and update the program counter.
[0044] In some embodiments, the adaptive prediction module is specifically used for: Predict the jump probability based on historical data and the jump prediction message to obtain the predicted probability; Predict the instruction quantity based on historical data and the predicted probability to obtain the predicted quantity; Determine the second instruction to be executed based on the predicted quantity and the jump address in the jump prediction message.
[0045] In some embodiments, the prediction probability can be determined based on the following formula : ; where , is a feature vector including historical jump frequency, jump distance, and instruction type; are model parameters estimated using training data, j used to distinguish different jump instructions, n represents the number of feature vectors.
[0046] As an example, the model parameters β can be trained and optimized through the following scheme.
[0047] Collect relevant data on historical jump instructions, including the address of the jump instruction, the jump target address, jump frequency, jump distance, etc.
[0048] Collect information on the type of instructions, such as branch instructions, jump instructions, call instructions, etc.
[0049] Collect information on the execution context of instructions, such as the execution order of instructions, the dependency relationship of instructions, etc.
[0050] Clean the collected data to remove outliers and noise.
[0051] Normalize the data to convert data of different scales to the same scale range.
[0052] Divide the data into a training set, a validation set, and a test set for training and evaluating the model.
[0053] Select features related to jump prediction, such as historical jump frequency, jump distance, instruction type, instruction execution context, etc.
[0054] Use feature selection techniques (such as principal component analysis PCA, Lasso regression, etc.) to reduce the dimensionality and select features, and extract the most predictive features.
[0055] Select a logistic regression model as the base model for predicting the execution probability of jump instructions.
[0056] The output of the logistic regression model is a probability value indicating the probability that the jump instruction is executed.
[0057] Use the training set data to train the logistic regression model and optimize the model parameters β .
[0058] The optimization objective is to maximize the likelihood function, that is, to maximize the prediction probability of the training data.
[0059] Use gradient descent or other optimization algorithms (such as Newton's method, quasi-Newton method, etc.) for parameter optimization.
[0060] The loss function of the logistic regression model is the cross-entropy loss function: ; where N is the number of training samples, y i is the actual jump label (1 means jump, 0 means no jump), p i is the predicted jump probability.
[0061] Use gradient descent for parameter optimization.
[0062] Use the validation set data to evaluate the trained model, and calculate metrics such as the prediction accuracy, recall rate, and F1 score of the model.
[0063] Adjust the model parameters and feature selection according to the evaluation results to optimize the model performance.
[0064] In practical applications, regularly update the model with new data to adapt to system changes and new jump patterns.
[0065] Use online learning techniques to update the model parameters in real time to improve the adaptability and prediction accuracy of the model.
[0066] As an example, the predicted quantity can be determined based on the following formula : ; where is the minimum instruction length of the cache; is the maximum instruction length of the cache.
[0067] As another example, the predicted quantity is determined based on the following formula : ; where is the minimum instruction length of the cache; is the maximum instruction length of the cache; δ is the adjustment coefficient used to dynamically adjust the instruction length according to the value of the feature vector X ; α1 ,α 2 ,⋯,α n is the eigenvector X and the weight coefficient of; i is used to distinguish different eigenvectors, n indicating the number of eigenvectors.
[0068] In the embodiment of the present application, the multi-instruction stream in the buffer is dynamically managed through label indication. When jumping, there is no need to clear the buffer or wait for the main memory to be reloaded. By using the atomic switching mechanism of the jump identifier and the activation status identifier (such as the state flip of the first label indication and the second label indication), the instruction stream switching at the nanosecond level is realized.
[0069] The label multiplexing technology (triggering the storage area coverage when the quantity identifier is reset to zero) is adopted to enable a single buffer to support the residence of multiple instruction streams, reducing the storage unit occupancy by 33% compared with the traditional double-buffer scheme; the preloaded instruction quantity is accurately controlled through an adaptive prediction model (jump probability calculation based on logistic regression and instruction quantity prediction formula), reducing the redundant data transfer by 18%.
[0070] The illegal jump is blocked through a dual label verification mechanism (the pre-decoder verifies the instruction format + the secondary decoder detects the memory permission); the design of the abnormal fallback guarantee (retaining the unexhausted part of the original instruction stream) shortens the error prediction recovery delay to 5 clock cycles, and the failure recovery success rate is increased to 99.7%.
[0071] It does not constitute a limitation on the technical solution provided by the embodiment of the present application. Those skilled in the art know that with the evolution of SOC technology and the emergence of new application scenarios, the technical solution provided by the embodiment of the present application is equally applicable to similar technical problems.
[0072] Next, refer to Figures 1 to 2 to describe the system-level chip caching method using the adaptive prediction technology according to the first aspect embodiment of the present invention.
[0073] As Figure 3 shown, the embodiment of the present application also provides a schematic flow diagram of a system-level chip caching method using the adaptive prediction technology, which specifically includes the following steps: S310, according to a preset rule, read a plurality of consecutive first to-be-executed instructions from the main memory and cache them into the buffer. The buffer also includes a first label indication for indicating the number of remaining unexecuted instructions in the first to-be-executed instructions; S320, based on the first label indication, read the first target to-be-executed instruction in the first to-be-executed instructions from the buffer to the decoding unit. The decoding unit determines the instruction type of the first target to-be-executed instruction, and the instruction type includes a jump type and a non-jump type; S330. When it is determined that the instruction type of the first target instruction to be executed is a jump type, send a jump prediction message to the adaptive prediction module; the adaptive prediction module makes a prediction based on the jump prediction message and sends the prediction result to the address selection module. S340. The address selection module reads multiple consecutive second instructions to be executed from the main memory and caches them in the buffer. The second instructions to be executed are the instructions to be executed after the jump. The buffer also includes a second label indication for indicating the number of remaining unexecuted instructions in the second instructions to be executed. S350. When the decoding unit determines to execute the first target instruction to be executed, send a jump execution message to the buffer; after receiving the jump execution message, the buffer reads the second target instruction to be executed in the second instructions to be executed to the decoder based on the second label indication.
[0074] In an embodiment of the present application, the main memory is responsible for reading a series of consecutive instructions and storing them in the buffer. The buffer is used to temporarily store these instructions and includes a label indication for indicating the number of unexecuted instructions. When the decoding unit receives these instructions, it identifies the type of the instructions and determines whether they are jump instructions. If they are jump instructions, the decoding unit sends jump prediction information to the adaptive prediction module, which makes a jump decision based on the prediction information and passes the result to the address selection module. The address selection module guides the main memory to read the instruction sequence after the jump according to the prediction result and stores it in the buffer. The buffer also includes another label indication for indicating the number of unexecuted instructions in the instruction sequence after the jump. When the decoding unit determines to execute the jump instruction, it sends a jump execution message to the buffer, and the buffer then reads the target instruction after the jump to the decoding unit according to this information. This embodiment manages the jump of instructions by using label indications, avoiding the need to empty the buffer or wait for the main memory to reload instructions when executing jump instructions, thereby significantly reducing the waiting time, improving the system efficiency, reducing resource waste, and improving the resource utilization rate.
[0075] In some embodiments, the buffer includes multiple labels, each label corresponding to a storage area, each area being used to store the first instructions to be executed or the second instructions to be executed read at the same time. Each label respectively includes: a quantity identifier, an activation status identifier, and a jump identifier. The label with the first state of the jump identifier corresponds to the first instruction to be executed, and the label with the second state of the jump identifier corresponds to the second instruction to be executed. The quantity identifier is used to record the number of unexecuted instructions in the corresponding area. When the number of unexecuted instructions corresponding to the label is 0, the label is used to record that the corresponding area can be overwritten. The activation status flag is used to indicate whether the to-be-executed instructions stored in the corresponding area are readable; Among them, the first tag indicates that the tag corresponding to the jump identifier in the first state is in the active state, and the tag corresponding to the jump identifier in the second state is in the non-active state; The second tag indicates that the tag corresponding to the jump identifier in the first state is in the non-active state, and the tag corresponding to the jump identifier in the second state is in the active state.
[0076] In some embodiments, the buffer includes multiple tags, and each tag is used to record a to-be-executed instruction; Among them, the first tag is used to record the first unexecuted to-be-executed instruction in the first to-be-executed instruction; The second tag is used to record the first unexecuted to-be-executed instruction in the second to-be-executed instruction; The third tag is used to record whether it is the first to-be-executed instruction or the second to-be-executed instruction currently being executed; Among them, the first tag indicates that the corresponding third tag indicates the execution of the first to-be-executed instruction; The second tag indicates that the corresponding third tag indicates the execution of the second to-be-executed instruction.
[0077] In some embodiments, the decoding unit includes a pre-decoder, a first-stage decoder, and a second-stage decoder connected in sequence; Among them, the pre-decoder confirms the instruction type of the first target to-be-executed instruction read; If the first target to-be-executed instruction is confirmed to be a jump instruction, a jump prediction message is generated; The first-stage decoder confirms the first jump requirement status of the first target to-be-executed instruction, and the first jump requirement status includes a first jump state and a first non-jump state; If the first jump requirement status is confirmed to be the first jump state, a jump execution message is generated.
[0078] In some embodiments, the adaptive prediction module makes predictions based on the jump prediction message, including: Predicting the jump probability based on historical data and the jump prediction message to obtain a prediction probability; Predicting the number of instructions based on historical data and the prediction probability to obtain a predicted number; Determining the second to-be-executed instruction based on the predicted number and the jump address in the jump prediction message.
[0079] Among them, the prediction probability and the predicted number can be combined with the descriptions of the solutions in the foregoing Figure 1 and Figure 2 The embodiments shown are not described in detail here.
[0080] The architecture platform described in the embodiments of this application is to more clearly illustrate the technical solutions of the embodiments of this application. In all the examples shown and described here, any specific value should be construed as merely exemplary and not as a limitation. Therefore, other examples of the exemplary embodiments may have different values.
[0081] It should be noted that like reference numerals and letters denote like items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0082] In addition, in the description of the embodiments of the present invention, unless otherwise clearly specified and limited, the terms "mounted", "connected", and "coupled" should be construed broadly. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the internal communication of two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood in specific situations.
[0083] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and should not be construed as indicating or implying relative importance.
[0084] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or equivalently replace some or all of the technical features. However, such modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A system-level chip cache method using adaptive prediction technology, characterized in that: The following steps are involved: According to a preset rule, a plurality of consecutive first instructions to be executed are read from a main memory and cached in a buffer, wherein the buffer also includes a first tag indication, and the first tag indication is used to indicate the number of remaining unexecuted instructions in the first instructions to be executed; Based on the first tag indication, a first target instruction to be executed in the first instruction to be executed is read from the buffer to a decoding unit, and the decoding unit determines an instruction type of the first target instruction to be executed, where the instruction type includes a jump type and a non-jump type; When it is determined that the instruction type of the first target to-be-executed instruction is a jump type, sending a jump prediction message to the adaptive prediction module; The adaptive prediction module makes predictions based on the jump prediction message and sends the prediction results to the address selection module; The address selection module reads a plurality of consecutive second instructions to be executed from the main memory based on the prediction result and caches them in a buffer, wherein the second instructions to be executed are instructions to be executed after a jump, and the buffer further includes a second label indication, wherein the second label indication is used to indicate the number of remaining unexecuted instructions in the second instructions to be executed; When the decoding unit determines to execute the first target instruction to be executed, a jump execution message is sent to the buffer; After receiving the jump execution message, the buffer reads a second target to-be-executed instruction in the second to-be-executed instructions to a decoder based on the second tag instruction.
2. The method according to claim 1, characterized in that The buffer includes a plurality of tags, each tag corresponds to a storage area, each area is used to store the first to-be-executed instruction or the second to-be-executed instruction read at the same time, and each tag includes: a quantity mark, an activation state mark and a jump mark; The label of the first state of the jump identifier corresponds to the first instruction to be executed, and the label of the second state of the jump identifier corresponds to the second instruction to be executed; The quantity identifier is used to record the number of unexecuted instructions in the corresponding area. When the number of unexecuted instructions corresponding to the tag is 0, the tag is used to record that the corresponding area can be overwritten; The activation status flag is used to indicate whether the to-be-executed instructions stored in the corresponding area can be read; The first label indicates that the label corresponding to the jump mark in the first state is in an activated state, and the label corresponding to the jump mark in the second state is in an inactivated state; The second label indicates that the label corresponding to the jump flag being in the first state is in an inactive state, and the label corresponding to the jump flag being in the second state is in an active state.
3. The method according to claim 1, characterized in that The buffer includes a plurality of tags, each tag is used to record a to-be-executed instruction; The first tag is used to record the first unexecuted to-be-executed instruction in the first to-be-executed instructions; The second tag is used to record the first unexecuted pending instruction in the second pending instructions; The third tag is used to record whether the currently executed first instruction to be executed is the second instruction to be executed; Wherein, the first tag indicates execution of the first instruction to be executed corresponding to the third tag instruction; The second tag indicates that the second to-be-executed instruction is executed corresponding to the third tag.
4. The method according to claim 1, characterized in that The decoding unit includes a pre-decoder and a multi-stage decoder connected in sequence; The pre-decoder confirms the instruction type of the first target instruction to be executed read; If the first target to-be-executed instruction is confirmed to be a jump instruction, generating a jump prediction message; Confirming, by the multi-stage decoder, a first jump requirement state of the first target instruction to be executed, wherein the first jump requirement state includes a first jump state and a first non-jump state; If the first jump requirement state is confirmed as the first jump state, a jump execution message is generated.
5. The method according to claim 1, characterized in that The adaptive prediction module makes predictions based on the jump prediction messages, including: The jump probability is predicted based on historical data and jump prediction messages to obtain the predicted probability; Predict the number of instructions based on historical data and prediction probability to obtain the predicted number; A second instruction to be executed is determined based on the predicted number and the jump address in the jump prediction message.
6. The method according to claim 5, characterized in that The predicted probability is determined based on the following formula : ; in, , including the feature vectors of historical jump frequency, jump distance, and instruction type; is the model parameter, estimated through training data, j is used to distinguish different jump instructions, and n represents the number of feature vectors.
7. The method according to claim 6, characterized in that The forecast quantity is determined based on the following formula : ; in, is the minimum instruction length of the cache; The maximum instruction length of the cache.
8. The method according to claim 6, characterized in that The forecast quantity is determined based on the following formula : ; in, is the minimum instruction length of the cache; is the maximum instruction length of the cache; δ is the adjustment coefficient, used to adjust the eigenvector X The value of dynamically adjusts the instruction length; α 1 ,α 2 ,⋯,α n is the feature vector X The weight coefficient of i Used to distinguish different feature vectors, n Represents the number of eigenvectors.
9. A system-on-chip cache system using adaptive prediction technology, characterized in that: include: Main memory, buffer, decoding unit, adaptive prediction module, address selection module; The main memory is used to read a plurality of consecutive first instructions to be executed according to a preset rule and cache them in a buffer, wherein the buffer also includes a first tag indication, and the first tag indication is used to indicate the number of remaining unexecuted instructions in the first instructions to be executed; The buffer is used to read a first target to-be-executed instruction in the first to-be-executed instructions to a decoding unit based on the first tag indication; The decoding unit is used to determine the instruction type of the first target instruction to be executed, where the instruction type includes a jump type and a non-jump type; When it is determined that the instruction type of the first target to-be-executed instruction is a jump type, sending a jump prediction message to the adaptive prediction module; The adaptive prediction module is used to make predictions based on the jump prediction message and send the prediction results to the address selection module; The address selection module is used to instruct the main memory to read a plurality of consecutive second instructions to be executed and cache them in a buffer based on the prediction result, wherein the second instructions to be executed are instructions to be executed after a jump, and the buffer also includes a second label indication, and the second label indication is used to indicate the number of remaining unexecuted instructions in the second instructions to be executed; The decoding unit is further configured to, when determining to execute the first target instruction to be executed, send a jump execution message to the buffer; The buffer is further configured to read a second target to-be-executed instruction in the second to-be-executed instruction to a decoder based on the second tag indication after receiving the jump execution message.
10. The system according to claim 9, characterized in that The buffer includes a plurality of tags, each tag corresponds to a storage area, each area is used to store the first to-be-executed instruction or the second to-be-executed instruction read at the same time, and each tag includes: a quantity mark, an activation state mark and a jump mark; The label of the first state of the jump identifier corresponds to the first instruction to be executed, and the label of the second state of the jump identifier corresponds to the second instruction to be executed; The quantity identifier is used to record the number of unexecuted instructions in the corresponding area. When the number of unexecuted instructions corresponding to the tag is 0, the tag is used to record that the corresponding area can be overwritten; The activation status flag is used to indicate whether the to-be-executed instructions stored in the corresponding area can be read; The first label indicates that the label corresponding to the jump mark in the first state is in an activated state, and the label corresponding to the jump mark in the second state is in an inactivated state; The second label indicates that the label corresponding to the jump flag being in the first state is in an inactive state, and the label corresponding to the jump flag being in the second state is in an active state.
Citation Information
Patent Citations
Hybrid branch prediction device with sparse and dense prediction caches
CN102160033A
Device and method for realizing indirect branch and prediction among modern processors
CN102306094A
Parallel branch prediction device of packet-based updating historical information
CN102520913A
Road prediction method for instruction cache, access control unit, and instruction processing device
CN112559049A
Caching method and system for SOC and SOC
CN114138335A