Data forwarding device and method, chip and computer equipment

By designing a data pre-delivery device in the processor, using the coordinated work of storage queues, storage buffers and data caches, the data loading delay problem of the processor in address dependence scenarios is solved, and a more efficient data loading process is achieved.

CN119938140APending Publication Date: 2025-05-06TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311454112.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-02
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

During the processor instruction process, when the current sequence data storage instruction and subsequent data loading instruction generate address dependence, the processor needs to identify the address dependence relationship to ensure that the data read by the data loading instruction is the latest data stored in the data storage instruction, resulting in an increase in delay during the data loading process.

Method used

A data pre-delivery device is designed, including a storage queue, a storage buffer, a data cache, a first pre-delivery merging unit and a second pre-delivery merging unit. Through the collaborative work of these components, the pre-passed data in the storage queue and the storage buffer is merged, and further merged with the data in the data cache to generate the latest pre-passed data.

Benefits of technology

Through the use of the data pre-passing device, the delay in reading data by data loading instructions can be significantly reduced, and the performance and efficiency of the processor can be improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938140A_ABST
    Figure CN119938140A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a data forwarding device and method, a chip and computer equipment, and belongs to the technical field of chips. The data forwarding device comprises a storage queue, a storage buffer, a data buffer, a first forwarding merging unit and a second forwarding merging unit, the first forward merging unit is used for reading the first forward data from the storage queue, reading the second forward data from the storage buffer, and performing forward merging on the first forward data and the second forward data to obtain third forward data; the second forward merging unit is used for reading the cached data from the data cache based on the data loading instruction, reading the third forward data from the first forward merging unit, and performing forward merging on the third forward data and the cached data to obtain fourth forward data; and the second forward merging unit is also used for outputting fourth forward data corresponding to the data loading instruction to the data loading pipeline. According to the embodiment of the invention, the data loading efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of chip technology, and in particular to a data forwarding device, method, chip and computer equipment. Background Art

[0002] During the processor instruction process, when a preceding data storage (store) instruction and a subsequent data loading (load) instruction generate an address dependency, the processor needs to identify the address dependency and ensure that the data read by the data loading instruction is the latest data stored by the data storage instruction based on the address dependency.

[0003] Data forwarding refers to the process of passing the execution result of the previous instruction directly to the subsequent instruction. In address-dependent scenarios, data forwarding can reduce the delay caused by data writeback and reading during data loading. Therefore, how to improve the performance of data forwarding has become an important issue in improving processor performance. Summary of the invention

[0004] The embodiment of the present application provides a data forwarding device, method, chip and computer equipment. The technical solution is as follows:

[0005] On the one hand, an embodiment of the present application provides a data forwarding device, the device comprising:

[0006] A storage queue, a storage buffer, a data cache, a first forward merging unit, and a second forward merging unit;

[0007] The storage queue and the storage buffer are respectively connected to the first forward merging unit, and the second forward merging unit is respectively connected to the data cache and the first forward merging unit;

[0008] The first forward merging unit is used to read the first forward data from the storage queue, read the second forward data from the storage buffer, and perform forward merging on the first forward data and the second forward data to obtain third forward data;

[0009] The second forward merging unit is used to read cache data from the data cache based on the data load instruction, read the third forward data from the first forward merging unit, and forward merge the third forward data and the cache data to obtain fourth forward data;

[0010] The second forward merge unit is further configured to output the fourth forward data corresponding to the data loading instruction to the data loading pipeline.

[0011] On the other hand, an embodiment of the present application provides a data forwarding method, the method comprising:

[0012] Reading the first forward data from the storage queue and the second forward data from the storage buffer by the first forward merging unit, and forward merging the first forward data and the second forward data to obtain the third forward data;

[0013] Based on the data load instruction, read the cache data from the data cache through the second forward merging unit, read the third forward data from the first forward merging unit, and forward merge the third forward data and the cache data to obtain fourth forward data;

[0014] The fourth forward data corresponding to the data loading instruction is output to the data loading pipeline through the second forward merging unit.

[0015] In some embodiments, forward merging the first forward data and the second forward data to obtain the third forward data includes:

[0016] When the first forward data contains data corresponding to the first address and the second forward data does not contain data corresponding to the first address, the data corresponding to the first address in the first forward data is selected as the third forward data; when the first forward data does not contain data corresponding to the second address and the second forward data contains data corresponding to the second address, the data corresponding to the second address in the second forward data is selected as the third forward data.

[0017] In some embodiments, the step of forward merging the third forward data and the cached data to obtain the fourth forward data includes:

[0018] In a case where the third forward data contains data corresponding to the third address and the cache data does not contain data corresponding to the third address, the data corresponding to the third address in the third forward data is selected as the fourth forward data; in a case where the third forward data does not contain data corresponding to the fourth address and the cache data contains data corresponding to the fourth address, the data corresponding to the fourth address in the cache data is selected as the fourth forward data.

[0019] In some embodiments, the forward delivery priority of the data in the storage queue is higher than the forward delivery priority of the data in the storage buffer, and the forward delivery priority of the data in the storage queue and the storage buffer is higher than the forward delivery priority of the data in the data cache;

[0020] The step of forward merging the first forward data and the second forward data to obtain the third forward data further includes:

[0021] In the case where there is an address overlap between the first forward data and the second forward data, selecting data belonging to the first forward data as the third forward data based on the overlapping address;

[0022] The forward merging of the third forward data and the cached data to obtain the fourth forward data also includes:

[0023] The second forward merging unit is further configured to select data belonging to the third forward data as the fourth forward data based on the overlapping addresses when there is address overlap between the third forward data and the cache data.

[0024] In some embodiments, the storage queue is provided with a dependent address determination unit and a forward data selection unit;

[0025] The method further comprises:

[0026] Based on the data loading instruction and the data storage instructions in the storage queue, determining a dependent address sequence by the dependent address determining unit, wherein the dependent address sequence is used to indicate instruction addresses of the data storage instructions on which the data loading instruction depends;

[0027] Based on the dependent address sequence, the first forward data is selected from the storage queue by the forward data selection unit.

[0028] In some embodiments, the dependency address determination unit includes a search range generator, a dependency relationship identifier, and a first AND operator;

[0029] The determining, by the dependent address determining unit, a dependent address sequence based on the data loading instruction and the data storage instruction in the storage queue comprises:

[0030] Determining a search range in the storage queue by the search range generator based on a head pointer position of the storage queue and a load instruction position of the data load instruction, wherein the load instruction position is used to represent an instruction position of the data load instruction in a data storage instruction sequence flow;

[0031] Based on the instruction address of the data storage instruction and the instruction address of the data loading instruction in the storage queue, determining an instruction dependency bitmap by the dependency identifier, wherein the instruction dependency bitmap is used to characterize the dependency relationship between the data loading instruction and each of the data storage instructions;

[0032] The dependency address sequence is obtained by performing a bitwise AND operation on the search range and the instruction dependency bitmap through the first AND operator.

[0033] In some embodiments, the storage queue is a circular queue;

[0034] The determining, by the search range generator, of the search range in the storage queue based on the head pointer position of the storage queue and the load instruction position of the data load instruction comprises:

[0035] In a case where the flag bit corresponding to the head pointer position is the same as the flag bit corresponding to the load instruction position, determining the range after the head pointer position and before the load instruction position in the storage queue as the search range;

[0036] When the flag bit corresponding to the head pointer position is different from the flag bit corresponding to the load instruction position, the storage queue is rewound; and a range after the head pointer position and before the load instruction position in the storage queue after the rewinding is determined as the search range;

[0037] The flag bit is 0 or 1, and the flag bit switches when it moves from the end of the team to the head of the team.

[0038] In some embodiments, the dependency identifier includes a comparator array and a replicator, the number of comparators in the comparator array being the same as the number of entries in the storage queue;

[0039] The step of determining the instruction dependency bitmap by the dependency identifier based on the instruction address of the data storage instruction and the instruction address of the data loading instruction in the storage queue comprises:

[0040] The instruction address of the data storage instruction is compared with the instruction address of each of the data loading instructions one by one by the comparator array to obtain an original bitmap, wherein the original bitmap is composed of 0 and 1, wherein 0 is used to indicate that the instruction address of the data storage instruction is inconsistent with the instruction address of the data loading instruction, and 1 is used to indicate that the instruction address of the data storage instruction is consistent with the instruction address of the data loading instruction;

[0041] The original bitmap is copied by the replicator to obtain the instruction dependency bitmap, and the length of the instruction dependency bitmap is twice that of the original bitmap.

[0042] In some embodiments, each entry in the storage queue is n bytes, the forward data selection unit is composed of n forward data selection subunits, and different forward data selection subunits correspond to different bytes in the entry;

[0043] The n forward data selection subunits operate in parallel, and the first forward data is obtained by merging data output by each of the forward data selection subunits.

[0044] In some embodiments, for the i-th forward data selection subunit among the n forward data selection subunits, the i-th forward data selection subunit includes a second AND operator and a tree selector;

[0045] The selecting the first forward data from the storage queue by the forward data selection unit based on the dependent address sequence includes:

[0046] Performing a bitwise AND operation on the dependent address sequence and the i-th byte valid sequence by the second AND operator to obtain an i-th byte valid dependent address sequence, wherein the i-th byte valid sequence is used to indicate the validity of the data stored in the i-th byte in each entry;

[0047] Based on the i-th byte effective dependent address sequence, the i-th forward byte is selected from the i-th bytes of each entry in the storage queue by the tree selector.

[0048] In some embodiments, based on the i-th byte effective dependent address sequence, selecting the i-th forward byte from the i-th byte of each entry in the storage queue by the tree selector comprises:

[0049] Determining a valid i-th byte in each entry in the storage queue based on the i-th byte valid dependent address sequence;

[0050] Based on the merging rule, the valid i-th bytes in the adjacent entries are merged through the first-level tree structure, and the merging result is input into the second-level tree structure;

[0051] Based on the merging rule, adjacent merging results outputted from the j-1th layer of tree structure are merged through the jth layer of tree structure, and the merging results are inputted into the j+1th layer of tree structure, where j is an integer greater than or equal to 2.

[0052] In some embodiments, the merging rule indicates that a merging priority of the i-th byte in the k-th entry is higher than a merging priority of the i-th byte in the k-1-th entry.

[0053] In some embodiments, the i-th forward data selection subunit further includes a first replicator and a second replicator;

[0054] The method further comprises:

[0055] Copying the i-th byte valid sequence by the first replicator, wherein the length of the i-th byte valid sequence after copying is consistent with the length of the dependent address sequence;

[0056] The storage queue is replicated by the second replicator, wherein the length of the storage queue after replication is consistent with the length of the dependent address sequence.

[0057] On the other hand, an embodiment of the present application provides a chip, which includes the data forwarding device as described in the above aspects.

[0058] On the other hand, an embodiment of the present application provides a computer device, which includes a processor and a memory, wherein the processor is connected to the memory via a bus, and the processor is provided with a data forwarding device as described in the above aspect.

[0059] In an embodiment of the present application, a first forward merging unit and a second forward merging unit are provided in a data forward delivery device, the first forward merging unit is used to merge the forward data in the storage queue and the storage buffer, and the second forward merging unit is used to further forward merge the cache data in the data cache and the merge result of the first forward merging unit, so that the data loading pipeline can obtain the latest storage data from the storage queue, the storage buffer and the data cache, and then directly feed back the latest storage data to the data loading instruction, thereby realizing data forward delivery from storage to loading, which helps to improve the execution efficiency of the data loading instruction, and thereby improve the performance of the chip configured with the data forward delivery device. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0061] Figure 1 A schematic diagram showing the structure of a data forwarding device provided by an exemplary embodiment of the present application is shown;

[0062] Figure 2 It is a schematic diagram of an implementation of a forward data merging process shown in an exemplary embodiment of the present application;

[0063] Figure 3 is a schematic diagram of a forward pass micro-architecture in a storage queue shown in an exemplary embodiment of the present application;

[0064] Figure 4 is a schematic diagram of the structure of an address dependency determination unit shown in an exemplary embodiment of the present application;

[0065] Figure 5 is a structural diagram of an address dependency determination unit shown in another exemplary embodiment of the present application;

[0066] Figure 6 is a schematic diagram of an implementation of a search scope determination process shown in an exemplary embodiment of the present application;

[0067] Figure 7 is a structural diagram of a forward data selection unit shown in an exemplary embodiment of the present application;

[0068] Figure 8 is a schematic diagram of an implementation of a forward data selection process shown in an exemplary embodiment of the present application;

[0069] Fig. 9 is a schematic diagram of a forward pass micro-architecture in a storage queue shown in another exemplary embodiment of the application;

[0070] Fig.10 A flow chart of a data forwarding method provided by an exemplary embodiment of the present application is shown;

[0071] Fig.11 A structural block diagram of a computer device provided by an exemplary embodiment of the present application is shown. DETAILED DESCRIPTION

[0072] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0073] It should be understood that the "several" mentioned in this article refers to one or more, and "multiple" refers to two or more. "And / or" describes the association relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the associated objects before and after are in an "or" relationship.

[0074] Please refer to Figure 1 , which shows a schematic diagram of the structure of a data forwarding device provided by an exemplary embodiment of the present application. The data forwarding device 100 includes a storage queue 110, a storage buffer 120, a data cache 130, a first forward merging unit 140 and a second forward merging unit 150.

[0075] The storage queue 110 and the storage buffer 120 are respectively connected to the first forward merging unit 140 , and the second forward merging unit 150 is respectively connected to the data cache 130 and the first forward merging unit 140 .

[0076] The store queue 110 is a component for temporarily storing a data storage instruction stream. Optionally, the store queue 110 can be implemented using a register or an SRAM (Static Random-Access Memory).

[0077] In some embodiments, the data storage instruction stream stored in the storage queue 110 is a sequential instruction stream.

[0078] In a possible implementation, after the data storage instruction passes through the out-of-order execution unit (not shown in the figure), it is written into the storage queue 110 for temporary storage, and the storage queue 110 is used to convert the out-of-order instruction into a sequential instruction stream according to the instruction sequence. The instruction sequence is the execution order of the instructions, that is, the execution order of the instructions indicated by the code.

[0079] For example, the order of data storage instructions output by the out-of-order execution unit is storage instruction 3, storage instruction 1 and storage instruction 2. After conversion based on the instruction execution order, the instructions stored in the storage queue 110 are storage instruction 1, storage instruction 2 and storage instruction 3 respectively.

[0080] In some embodiments, the storage queue 110 includes a plurality of entries for storing data storage instructions. In an illustrative example, each entry in the storage queue 110 is 64 bits (8 bytes), that is, each entry can store 64 bits of data.

[0081] The store buffer 120 is a component located after the store queue 110 and used for buffering data storage instructions. Optionally, the store buffer 120 can be implemented by a register or SRAM.

[0082] In some embodiments, the data storage instructions in the storage queue 110 are written into the storage buffer 120 according to the First Input First Output (FIFO) principle.

[0083] In some embodiments, the storage buffer 120 is used to eliminate the speed mismatch problem between instruction generation and instruction execution. In addition, the storage buffer 120 is also used to merge the storage data indicated by the data storage instruction to reduce the number of accesses to the lower-level cache.

[0084] In some embodiments, storage buffer 120 includes a plurality of entries for storing data storage instructions.

[0085] The data cache (Data Cache) 130 is a memory for caching data, and the memory can be implemented by SRAM or DRAM (Dynamic Random Access Memory).

[0086] In some embodiments, the data cache 130 stores frequently accessed data to improve the hit rate of the frequently accessed data. Optionally, the data cache 130 can improve the cache hit rate through various cache replacement algorithms.

[0087] In some embodiments, the data stored in the data cache 130 may come from the storage buffer 120 , that is, the storage buffer 120 is connected to the data cache 130 for writing the buffered data into the data cache 130 .

[0088] In some embodiments, when data is loaded, the data loading instruction will directly access the data cache 130 to read the data. Accordingly, the data cache 130 searches for corresponding data according to the address of the data to be loaded requested by the data loading instruction.

[0089] The first forward merging unit 140 and the second forward merging unit 150 are both elements for performing forward data merging.

[0090] The first forward merging unit 140 is used to read the first forward data from the storage queue 110 and the second forward data from the storage buffer 120, and forward merge the first forward data and the second forward data to obtain the third forward data.

[0091] In some embodiments, the storage queue 110 and the storage buffer 120 determine data having a storage address dependency relationship with the data loading instruction based on the data loading instruction, and pass the data as forward data to the first forward merging unit 140. Since there are differences between the data stored in the storage queue 110 and the storage buffer 120 (the data stored in the storage queue 110 is younger than the data stored in the storage buffer 120), the first forward merging unit 140 needs to merge the forward data after reading the forward data from the storage queue 110 and the storage buffer 120 to ensure that the third forward data obtained by merging is the latest data that has an address dependency relationship with the data loading instruction.

[0092] Optionally, the first forward merging unit 140 forward merges the first forward data and the second forward data based on the forward priorities corresponding to the storage queue 110 and the storage buffer 120 to obtain third forward data.

[0093] The second forward merging unit 150 is used to read cache data from the data cache 130 based on the data loading instruction, read third forward data from the first forward merging unit 140, and forward merge the third forward data and the cache data to obtain fourth forward data.

[0094] The data loading pipeline directly accesses the data cache 130 to obtain the hit cache data based on the data loading instruction. Since part of the data requested to be loaded by the data loading instruction may not be cached in the data cache 130, but is located in the storage queue 110 or the storage buffer 120, in order to avoid the loading delay caused by the write-back access of this part of the data, the second forward merging unit 150 further forward merges the cache data and the third forward data merged by the first forward merging unit 140 to obtain the fourth forward data.

[0095] Furthermore, the second forward merge unit 150 outputs the fourth forward data corresponding to the data loading instruction to the data loading pipeline, thereby completing the data forward transfer of the data loading instruction.

[0096] Optionally, the second forward merging unit 150 forward merges the third forward data and the cache data based on the forward priorities among the data cache 130 , the storage queue 110 , and the storage buffer 120 to obtain fourth forward data.

[0097] Since the data storage instruction may only hit part of the loading address range, for example, a data storage instruction instructs to store 1 byte, and the storage address is 2, when there is a subsequent data loading instruction that has an address dependency relationship with the data storage instruction, and the data loading instruction instructs to load 4 bytes, and the loading address range is 0 to 3, the problem of the data storage instruction hitting part of the loading address range occurs.

[0098] In order to quickly obtain the data required for the data loading instruction even in this partial hit scenario, in some embodiments, the first forward merge unit 140 and the second forward merge unit 150 are both byte-based forward merge units, that is, the first forward merge unit 140 and the second forward merge unit 150 perform forward merge in units of bytes.

[0099] To summarize, in the embodiments of the present application, a first forward merging unit and a second forward merging unit are provided in the data forward delivery device, the first forward merging unit is used to merge the forward data in the storage queue and the storage buffer, and the second forward merging unit is used to further forward merge the cache data in the data cache and the merging result of the first forward merging unit, so that the data loading pipeline can obtain the latest storage data from the storage queue, the storage buffer and the data cache, and then directly feed back the latest storage data to the data loading instruction, thereby realizing the forward delivery of data from storage to loading, which helps to improve the execution efficiency of the data loading instruction, and thereby improve the performance of the chip configured with the data forward delivery device.

[0100] In addition, the forward merge unit performs byte-level forward merge. When a data storage instruction hits a partial loading address range, it can also quickly merge the data required for the data loading instruction, further improving the data loading efficiency and thereby improving the performance of the chip configured with the data forward merge device.

[0101] Since the data requested to be loaded by the data loading instruction may come from the data cache, storage buffer and storage queue, and the age of the data stored in the above three is different, in order to ensure that the latest data is forwarded to the data loading instruction, the forward merge unit needs to perform forward merge based on the forward priority of the above three.

[0102] Since the data in the storage queue is younger than the data in the storage buffer, and the data in the storage buffer is younger than the data in the data cache, the forward priority of the above three is: storage queue>storage buffer>data cache. The forward merging process of the first and second forward merging units is described below in combination with the above forward priority.

[0103] In a possible implementation, when the first forward data includes data corresponding to the first address, and the second forward data does not include data corresponding to the first address, the first forward merging unit 140 selects data corresponding to the first address in the first forward data as the third forward data. When the first forward data does not include data corresponding to the second address, and the second forward data includes data corresponding to the second address, the first forward merging unit 140 selects data corresponding to the second address in the second forward data as the third forward data.

[0104] The first address and the second address are both addresses of data to be loaded requested by the data loading instruction.

[0105] When there is data corresponding to the first address in the storage queue 110, but there is no data corresponding to the first address in the storage buffer 120, it indicates that the data storage instruction corresponding to the first address has not been written into the storage buffer 120 by the storage queue 110. Therefore, the first forward merge unit 140 uses the data corresponding to the first address in the storage queue 110 as the merged forward data.

[0106] Similarly, when there is no data corresponding to the second address in the storage queue 110, but there is data corresponding to the second address in the storage buffer 120, it indicates that the data storage instruction corresponding to the second address has been written into the storage buffer 120 by the storage queue 110, so the first forward merge unit 140 uses the data corresponding to the second address in the storage buffer 120 as the merged forward data.

[0107] When there is an address overlap between the first forward data and the second forward data, the first forward merging unit 140 selects data belonging to the first forward data as the third forward data based on the overlapping address.

[0108] When there are data corresponding to the same address in the storage queue 110 and the storage buffer 120 (i.e., there is address overlap), it indicates that the storage queue 110 first stores the data storage instruction corresponding to the address, writes the data storage instruction into the storage buffer 120, and then stores a new data storage instruction corresponding to the address, and the new data storage instruction has not yet been written into the storage buffer 120. In this case, in order to ensure that the forward data is the latest storage data at the address, and because the forward priority of the storage queue 110 is higher than the forward priority of the storage buffer 120, the first forward merging unit 140 uses the data corresponding to the overlapping address in the storage queue 110 as the merged forward data.

[0109] In an illustrative example, Figure 2 As shown, the first forward data read from the storage queue by the first forward merging unit is valid for addresses 2, 4, and 7, and the second forward data read from the storage buffer is valid for addresses 1, 2, and 5. For addresses 4 and 7, since only the first forward data contains data of addresses 4 and 7, the data of addresses 4 and 7 in the merged third forward data come from the first forward data; for address 2, since both the first and second forward data contain data of address 2, based on the forward priority, the data of address 2 in the merged third forward data comes from the first forward data; for addresses 1 and 5, since only the second forward data contains data of addresses 1 and 5, the data of addresses 1 and 5 in the merged third forward data comes from the second forward data.

[0110] In the case where the third forward data includes data corresponding to the third address, and the cache data does not include data corresponding to the third address, the second forward merging unit 150 selects data corresponding to the third address in the third forward data as the fourth forward data. In the case where the third forward data does not include data corresponding to the fourth address, and the cache data includes data corresponding to the fourth address, the second forward merging unit 150 selects data corresponding to the fourth address in the cache data as the fourth forward data.

[0111] The third address and the fourth address are both addresses of data to be loaded requested by the data loading instruction.

[0112] When the third forward data contains data corresponding to the third address (from the storage queue or storage buffer), but the data corresponding to the third address does not exist in the data cache 130, it indicates that the data storage instruction corresponding to the third address has not yet been written into the data cache 130. Therefore, the second forward merging unit 150 uses the data corresponding to the third address in the third forward data as the merged forward data.

[0113] Similarly, when the data corresponding to the fourth address does not exist in the third forward data, but the data corresponding to the fourth address exists in the data cache 130, it indicates that the data requested to be loaded by the data loading instruction has been written into the cache and hit, and there is no data storage instruction that is address dependent on the data loading instruction in the current storage queue 110 and the storage buffer 120. Therefore, the second forward merge unit 150 uses the data corresponding to the fourth address in the storage data cache 130 as the merged forward data.

[0114] In the case where there is an address overlap between the third forward data and the cache data, the second forward merging unit 150 selects data belonging to the third forward data as the fourth forward data based on the overlapping address.

[0115] When the third forward data and the data cache 130 have data corresponding to the same address (i.e., there is address overlap), it indicates that the storage queue 110 or the storage buffer 120 stores the latest data corresponding to the overlapping address, while the data corresponding to the overlapping address in the data cache 130 is historical data. In this case, in order to ensure that the forward data is the latest stored data at the address, and because the forward priorities of the storage queue 110 and the storage buffer 120 are higher than the forward priority of the data cache 130, the second forward merging unit 150 uses the data corresponding to the overlapping address in the third forward data as the merged forward data.

[0116] In this embodiment, the forward merging unit forward merges the data obtained from the storage queue, storage buffer and data cache based on the forward priorities among the above three, ensuring that the forward data finally provided to the data load instruction is the latest storage data.

[0117] In some embodiments, the storage queue and storage buffer are both provided with a data forward micro-architecture, which is used to determine the forward data that needs to be provided to the forward merge unit based on the data load instruction. The following describes the process of determining the forward data in conjunction with the data forward micro-architecture in the storage queue.

[0118] Please refer to Figure 3 , which shows a schematic diagram of the structure of a data forward micro-architecture in a storage queue provided by an exemplary embodiment of the present application. The structure includes a dependent address determination unit 310 and a forward data selection unit 320.

[0119] The dependent address determining unit 310 is used to determine a dependent address sequence based on the data loading instruction and the data storage instructions in the storage queue, where the dependent address sequence is used to indicate the instruction addresses of the data storage instructions on which the data loading instruction depends.

[0120] In some embodiments, the input of the dependent address determination unit 310 includes a forward query request corresponding to the data loading instruction, and the forward query request includes the address of the data loading instruction. Based on the address of the data loading instruction and the addresses of each data storage instruction in the storage queue, the dependent address determination unit 310 determines the existing address dependency and determines the dependent address sequence.

[0121] Optionally, the addresses of the data loading instruction and the data storing instruction are virtual addresses (vaddr). The dependent address determining unit 310 determines whether there is an address dependency by comparing the virtual addresses.

[0122] In some embodiments, the dependent address sequence can be represented in a vector form (dep_vec), the vector is composed of 0 and 1, and the length of the vector is related to the number of entries in the storage queue. A 1 in the vector indicates that the data load instruction has a dependency relationship with the data storage instruction corresponding to the entry in the storage queue, and a 0 in the vector indicates that the data storage instruction corresponding to the entry in the data load instruction storage queue does not have a dependency relationship.

[0123] The forward data selection unit 320 is used to select the first forward data from the storage queue based on the dependent address sequence.

[0124] Based on the address dependency determined by the dependent address determining unit 310 , the forward data selecting unit 320 further combines the data stored in the storage queue to select the first forward data to be transmitted from the storage queue.

[0125] Regarding the detailed structure of the dependent address determination unit, such as Figure 4As shown, the dependent address determining unit 310 includes a search range generator 311 , a dependency relationship identifier 312 , and a first AND operator 313 .

[0126] The search range generator 311 is used to determine the search range in the storage queue based on the head pointer position of the storage queue and the load instruction position of the data load instruction, where the load instruction position is used to represent the instruction position of the data load instruction in the data storage instruction sequence flow.

[0127] Since there may be an address dependency relationship between the data loading instruction and the preceding data storage instruction, the search range generator 311 is required to determine the search range for which address dependency search is required from the storage queue.

[0128] In some embodiments, in order to determine which data storage instructions in the storage queue are located before the data loading instruction, the retrieval range generator 311 determines the head pointer position (Stq_deq_Idx) of the storage queue, and determines the loading instruction position (Forward.req.stqIdx) of the data loading instruction in the data storage instruction sequence stream, and determines the retrieval range with the above two positions.

[0129] Indicatively, Figure 5 As shown, the check range generator 311 outputs a check range based on the head pointer position and the load instruction position.

[0130] Optionally, both the head pointer position and the load instruction position are represented by entry indices of entries in the storage queue.

[0131] In an illustrative example, when the storage queue includes data storage instruction 1 , data storage instruction 2 , data storage instruction 3 and data storage instruction 4 , and the data loading instruction is located before the data storage instruction 4 , the loading instruction position is entry index 4 .

[0132] In a possible implementation, the storage queue adopts a circular queue. A circular queue is a linear data structure whose operation is based on the first-in-first-out principle and the tail of the queue is connected to the head of the queue to form a cycle.

[0133] In order to identify the wraparound of the load instruction position and thus improve the accuracy of the determined search range, in some embodiments, the head pointer position and the load instruction position are both corresponding to a flag, wherein the flag is 0 or 1, and the flag is switched when it moves from the end of the queue to the head of the queue.

[0134] In an illustrative example, the flag bits corresponding to the head pointer position are initially 0. Initially, the head pointer is at the head of the team and gradually moves to the end of the team. When the head pointer at the end of the team moves to the beginning of the team, the flag bit corresponding to the head pointer position changes from 0 to 1. When the head pointer gradually moves to the end of the team and then moves from the end of the team to the beginning of the team again, the flag bit corresponding to the head pointer position changes from 1 to 0.

[0135] In some embodiments, when determining whether a load instruction position wrap occurs, the search range generator detects whether the head pointer position is consistent with a flag bit corresponding to the load instruction position.

[0136] When the flag bit corresponding to the head pointer position is the same as the flag bit corresponding to the load instruction position, the search range generator 311 determines the range after the head pointer position and before the load instruction position in the storage queue as the search range.

[0137] When the flag bit corresponding to the head pointer position is the same as the flag bit corresponding to the load instruction position, it indicates that the load instruction position has not wrapped around, so the data storage instructions located after the head pointer position and before the load instruction position are all predecessor instructions of the data load instruction.

[0138] When the flag bit corresponding to the head pointer position is different from the flag bit corresponding to the load instruction position, the search range generator 311 wraps the storage queue and determines the range after the head pointer position and before the load instruction position in the storage queue after wrapping as the search range.

[0139] When the flag bit corresponding to the head pointer position is different from the flag bit corresponding to the load instruction position, it indicates that the load instruction position wraps around, that is, there is a preceding data storage instruction of the data load instruction before the load instruction position. Therefore, the retrieval range generator 311 also needs to wrap the storage queue to ensure that all preceding data storage instructions of the data load instruction in the storage queue are retrieved.

[0140] Optionally, wrapping the storage queue means: copying the storage queue and splicing it after the original storage queue. Accordingly, after wrapping the storage queue, the load instruction position needs to be moved to the copied storage queue.

[0141] After the storage queue wraps around, the instructions after the head pointer position and before the load instruction position are the predecessor instructions of the data load instruction.

[0142] In some embodiments, the search range may be represented by a bitmap consisting of 0 and 1. In addition, considering the wraparound situation, the length of the search range bitmap is twice the number of entries in the storage queue.

[0143] In an illustrative example, Figure 6As shown, the storage queue includes 4 entries (entry indexes are 0, 1, 2, and 3). When the head pointer position corresponds to entry index 0 and the load instruction position corresponds to entry index 2, since the flag bits corresponding to the head pointer position and the load instruction position are both 0, that is, no wraparound occurs, the search range is the entries with indexes 0 and 1, and the search range bitmap can be expressed as 11000000.

[0144] When the head pointer position corresponds to the entry index 3 and the load instruction position corresponds to the entry index 2, since the flag bits corresponding to the head pointer position and the load instruction position are different, a wraparound occurs. Therefore, after the storage queue is wrapped around, the retrieval range is determined to be entries with indexes 0, 1, and 3. The retrieval range bitmap can be represented as 00011100.

[0145] The dependency identifier 312 is used to determine the instruction dependency bitmap based on the instruction address of the data storage instruction in the storage queue (stq.vaddr, i.e., the address of the stored data) and the instruction address of the data loading instruction (forward.req.vaddr, i.e., the address of the data to be loaded). The instruction dependency bitmap is used to characterize the dependency between the data loading instruction and each data storage instruction.

[0146] Since the data loading instruction does not have an address dependency relationship with all the data storage instructions in the storage queue, it is necessary to use the dependency identifier 312 to determine the data storage instructions in the storage queue that have an address dependency relationship with the data loading instruction.

[0147] Among them, the data storage instructions that have an address dependency relationship with the data loading instructions in the storage queue can be represented by an instruction dependency bitmap. The instruction dependency bitmap is a bitmap composed of 0 and 1, and each element in the bitmap corresponds to an entry in the storage queue. If the element in the bitmap is 0, it means that there is no address dependency relationship between the data storage instruction stored in the entry corresponding to the element and the data loading instruction, and if the element in the bitmap is 1, it means that there is an address dependency relationship between the data storage instruction stored in the entry corresponding to the element and the data loading instruction.

[0148] In some embodiments, the dependency identifier 312 generates an instruction dependency bitmap by detecting whether the instruction address of the data loading instruction is consistent with the instruction address of each data storage instruction in the storage queue.

[0149] Detailed results on the dependency identifier, in one possible design, such as Figure 5 As shown, the dependency identifier 312 includes a comparator array 3121 and a replicator 3122 .

[0150] The number of comparators in the comparator array 3121 is the same as the number of entries in the storage queue. Figure 5 As shown, the number of entries in the storage queue is 4, and the comparator array 3121 includes 4 comparators.

[0151] The comparator array 3121 is used to compare the instruction address of the data storage instruction with the instruction address of each data loading instruction one by one to obtain an original bitmap (vaddr_match_bitmap). The original bitmap is composed of 0 and 1, 0 is used to indicate that the instruction address of the data storage instruction is inconsistent with the instruction address of the data loading instruction (that is, there is no address dependency), and 1 is used to indicate that the instruction address of the data storage instruction is consistent with the instruction address of the data loading instruction (that is, there is an address dependency).

[0152] like Figure 5 As shown, the input of each comparator in the comparator array 3121 includes the instruction address of the data loading instruction and the instruction address of the data storage instruction in an entry in the storage queue. When the addresses are consistent, the comparator outputs 1, and when the addresses are inconsistent, the comparator outputs 0. The output values ​​of each comparator are spliced ​​to obtain an original bitmap with a length of 4.

[0153] Due to the wraparound situation taken into account when determining the search range, the length of the search range bitmap determined is 2 times the number of entries. Therefore, in order to ensure that the length of the instruction dependency bitmap is consistent with that of the search range bitmap, the original bitmap is duplicated by the duplicator 3122 to obtain the instruction dependency bitmap, and the length of the instruction dependency bitmap is twice that of the original bitmap.

[0154] In an illustrative example, when the original bitmap is 1010, after passing through the copier, the output instruction dependency bitmap is 10101010.

[0155] Furthermore, the first AND operator 313 performs a bitwise AND operation on the search range and the instruction dependency bitmap, that is, addresses within the search range and having address dependencies are determined to obtain a dependency address sequence.

[0156] In an illustrative example, when the retrieval range bitmap is 11100000 and the instruction dependency bitmap is 10101010, the resulting dependency address sequence is 10100000, that is, the data storage instructions in the entries with indexes 0 and 2 in the storage queue have an address dependency with the data loading instructions and are located before the data loading instructions.

[0157] After the dependent address sequence is obtained by the address dependency determination unit 310, the dependent address sequence is further input into the forward data selection unit 320 for forward data selection.

[0158] Since the data requested to be loaded by the data load instruction may be part of the bytes in a certain entry, for example, each entry in the storage queue is 64 bits, and the data load instruction only needs to obtain the first byte of data in the entry, without the other 7 bytes of data.

[0159] Therefore, in some embodiments, the forward data selection unit 320 needs to select the forward data according to the validity of the data stored in the storage queue.

[0160] In one possible implementation, when each entry in the storage queue is n bytes, the forward data selection unit 320 can perform serial forward selection on each byte in the entry, that is, first select the data in the first byte that needs to be forwarded in each entry based on the dependency address sequence and the validity of the data of the first byte in each entry, and then select the data in the second byte that needs to be forwarded in each entry based on the dependency address sequence and the validity of the data of the second byte in each entry, and so on, until the data in the nth byte that needs to be forwarded in each entry is obtained.

[0161] To further improve the efficiency of forward data selection, in a possible design, when each entry in the storage queue is n bytes, the forward data selection unit 320 is composed of n forward data selection sub-units, and different forward data selection sub-units correspond to different bytes in the entry.

[0162] When performing forward data selection, n forward data selection subunits operate in parallel, and the first forward data is obtained by merging the data output by each forward data selection subunit. Since the n forward data selection subunits perform parallel forward selection (parallel forward), all forward data selection can be completed within one clock cycle (serial requires n clock cycles), thereby improving the selection efficiency of the forward data.

[0163] Indicatively, Figure 8 As shown, the forward data selection unit 320 is composed of 8 forward data selection sub-units 321, which are respectively used to select forward data from the 1st to the 8th bytes of the entry.

[0164] Regarding the detailed structure of each forward data selection subunit, in a possible design, such as Figure 8 As shown, each forward data selection subunit 321 includes a second AND operator 3211 and a tree selector 3212 .

[0165] In some embodiments, the input of each forward data selection subunit 321 includes two paths, one path is the dependent address sequence, and the other path is the byte valid sequence (stq_byte_valid). The dependent address sequence input to each forward data selection subunit 321 is consistent, while the byte valid sequence input to each forward data selection subunit 321 is different.

[0166] For the i-th forward data selection subunit, the input i-th byte valid sequence is used to indicate the validity of the data stored in the i-th byte in each entry of the storage queue.

[0167] Optionally, the i-th byte valid sequence is also represented by a bitmap of 0 and 1. For example, when the storage queue contains 4 entries and each entry is 8 bytes, the first byte valid sequence 0011 indicates that the data stored in the first byte of the 1st and 2nd entries are invalid, and the data stored in the first byte of the 3rd and 4th entries are valid; the third byte valid sequence 1010 indicates that the data stored in the third byte of the 1st and 3rd entries are valid, and the data stored in the third byte of the 2nd and 4th entries are invalid.

[0168] In some embodiments, the second AND operator 3211 performs a bitwise AND operation on the dependent address sequence and the i-th byte valid sequence to obtain the i-th byte valid dependent address sequence, which is used to indicate the validity of the data stored in the i-th byte in each entry.

[0169] In some embodiments, the i-th byte valid dependency address sequence is used to indicate whether the i-th byte in the entry with address dependency is valid. For example, the first byte valid dependency address sequence 10010000 indicates that the first byte of data and the data load instruction in the 1st and 4th entries are stored in the address dependency and are valid.

[0170] Since the dependent address sequence is copied, in order to ensure that the length of the i-th byte valid sequence is consistent with the length of the dependent address sequence, in some embodiments, the i-th forward data selection subunit also includes a first copier, which is used to copy the i-th byte valid sequence, wherein the length of the i-th byte valid sequence after copying is consistent with the length of the dependent address sequence.

[0171] For example, the original valid sequence of the i-th byte is 0011. After being copied by the first copier, the valid sequence of the i-th byte is 00110011.

[0172] The tree selector 3212 is used to select the i-th forward byte from the i-th byte of each entry in the storage queue based on the i-th byte effective dependent address sequence.

[0173] In some embodiments, the tree selector 3212 is composed of multiple layers of tree structures, and each layer of the tree structure is used to perform adjacent merging on the output of the previous layer of the tree structure, and finally determine the result of the output of the last layer of the tree structure as the i-th forward byte that needs to be forwarded in the i-th byte of each entry.

[0174] In order to match the length of the dependent address sequence, the forward data selection subunit 321 needs to perform a copy process on the storage queue.

[0175] In a possible design, the i-th forward data selection subunit further includes a second duplicator. The second duplicator is used to perform duplication processing on the storage queue, wherein the length of the storage queue after duplication is consistent with the length of the dependent address sequence.

[0176] When the length of the storage queue (i.e., the number of entries) is stq_depth, the length after the copy process is 2*stq_depth. Figure 7 As shown, the original length of the storage queue is 4, and the length of the storage queue after replication is 8.

[0177] In some embodiments, the tree selector 3212 is used to:

[0178] Determine the valid i-th byte in each entry in the storage queue based on the i-th byte valid dependent address sequence;

[0179] Based on the merging rule, the valid i-th bytes in the adjacent entries are merged through the first-level tree structure, and the merging result is input into the second-level tree structure;

[0180] Based on the merging rule, adjacent merging results outputted from the j-1th layer of the tree structure are merged through the jth layer of the tree structure, and the merging results are inputted into the j+1th layer of the tree structure, where j is an integer greater than or equal to 2.

[0181] like Figure 7 As shown, for a storage queue of length 8 (after replication), the tree selector 3212 performs level-by-level merging in the manner of 4→2→1.

[0182] Regarding the merge rule, in one possible implementation, the valid data to the right in the search range is younger. Accordingly, the merge rule indicates that the merge priority of the i-th byte in the k-th entry is higher than the merge priority of the i-th byte in the k-1-th entry, that is, the merge result of the i-th bytes in the k-th and k-1-th entries is the i-th byte in the k-th entry.

[0183] Of course, when the valid data on the left side of the search range is younger, the merge rule may indicate that the merge priority of the i-th byte in the k-th entry is lower than the merge priority of the i-th byte in the k-1-th entry, and this embodiment does not limit this.

[0184] In an illustrative example, Figure 8 As shown, when the byte effective dependent address sequence output by the second AND operator 3211 is 11010000, the data stored in the effective bytes are ABD respectively. The tree selector 3212 merges the effective bytes in the first and second entries through the first layer of the tree structure, and obtains the merged result as the effective byte in the second entry (i.e., B), and merges the effective bytes in the third and fourth entries, and obtains the merged result as the effective byte in the fourth entry (i.e., D). The remaining copied entries do not contain effective bytes, so the result is empty.

[0185] Furthermore, the tree selector 3212 merges the valid bytes in the second entry and the fourth entry through the second-level tree structure, and the merged result is the valid bytes in the fourth entry.

[0186] Finally, the tree selector 3212 outputs the valid byte in the 4th entry as the final forward byte through the third-level tree result.

[0187] Combination Figure 5 and Figure 7 As shown in the structure, when each storage data in the storage queue is 64 bits, the forward data microframe in the storage queue is as follows Fig. 9 shown.

[0188] In this embodiment, the forward data selection unit merges different bytes step by step through a multi-shard parallel merging strategy, which helps to shorten the merging efficiency of the forward data in the storage queue, thereby increasing the output speed of the forward data and improving the execution efficiency of the data loading instructions.

[0189] It should be noted that the forward data microstructure in the storage buffer can refer to the forward data microarchitecture of the storage queue in the above embodiment, and will not be described in detail in the embodiment of the present application.

[0190] Please refer to Fig.10 , which shows a flow chart of a data forwarding method provided by an exemplary embodiment of the present application, the method is used in the data forwarding management device provided by each of the above embodiments, and the method includes:

[0191] Step 1001: read first forward data from a storage queue and second forward data from a storage buffer through a first forward merging unit, and forward merge the first forward data and the second forward data to obtain third forward data.

[0192] Step 1002, based on the data loading instruction, read the cache data from the data cache through the second forward merging unit, read the third forward data from the first forward merging unit, and forward merge the third forward data and the cache data to obtain fourth forward data.

[0193] Step 1003: output fourth forward data corresponding to the data loading instruction to the data loading pipeline through the second forward merging unit.

[0194] In some embodiments, forward merging the first forward data and the second forward data to obtain the third forward data includes:

[0195] In the case where the first forward data contains data corresponding to the first address, and the second forward data does not contain data corresponding to the first address, the data corresponding to the first address in the first forward data is selected as the third forward data; in the case where the first forward data does not contain data corresponding to the second address, and the second forward data contains data corresponding to the second address, the data corresponding to the second address in the second forward data is selected as the third forward data;

[0196] In some embodiments, the step of forward merging the third forward data and the cached data to obtain the fourth forward data includes:

[0197] In a case where the third forward data contains data corresponding to the third address and the cache data does not contain data corresponding to the third address, the data corresponding to the third address in the third forward data is selected as the fourth forward data; in a case where the third forward data does not contain data corresponding to the fourth address and the cache data contains data corresponding to the fourth address, the data corresponding to the fourth address in the cache data is selected as the fourth forward data.

[0198] In some embodiments, the forward delivery priority of the data in the storage queue is higher than the forward delivery priority of the data in the storage buffer, and the forward delivery priority of the data in the storage queue and the storage buffer is higher than the forward delivery priority of the data in the data cache;

[0199] The step of forward merging the first forward data and the second forward data to obtain the third forward data further includes:

[0200] In the case where there is an address overlap between the first forward data and the second forward data, data belonging to the first forward data is selected as the third forward data based on the overlapping addresses.

[0201] The forward merging of the third forward data and the cached data to obtain the fourth forward data also includes:

[0202] The second forward merging unit is further configured to select data belonging to the third forward data as the fourth forward data based on the overlapping addresses when there is address overlap between the third forward data and the cache data.

[0203] In some embodiments, the storage queue is provided with a dependent address determination unit and a forward data selection unit;

[0204] The method further comprises:

[0205] Based on the data loading instruction and the data storage instructions in the storage queue, determining a dependent address sequence by the dependent address determining unit, wherein the dependent address sequence is used to indicate instruction addresses of the data storage instructions on which the data loading instruction depends;

[0206] Based on the dependent address sequence, the first forward data is selected from the storage queue by the forward data selection unit.

[0207] In some embodiments, the dependency address determination unit includes a search range generator, a dependency relationship identifier, and a first AND operator;

[0208] The determining, by the dependent address determining unit, a dependent address sequence based on the data loading instruction and the data storage instruction in the storage queue comprises:

[0209] Determining a search range in the storage queue by the search range generator based on a head pointer position of the storage queue and a load instruction position of the data load instruction, wherein the load instruction position is used to represent an instruction position of the data load instruction in a data storage instruction sequence flow;

[0210] Based on the instruction address of the data storage instruction and the instruction address of the data loading instruction in the storage queue, determining an instruction dependency bitmap by the dependency identifier, wherein the instruction dependency bitmap is used to characterize the dependency relationship between the data loading instruction and each of the data storage instructions;

[0211] The dependency address sequence is obtained by performing a bitwise AND operation on the search range and the instruction dependency bitmap through the first AND operator.

[0212] In some embodiments, the storage queue is a circular queue;

[0213] The determining, by the search range generator, of the search range in the storage queue based on the head pointer position of the storage queue and the load instruction position of the data load instruction comprises:

[0214] In a case where the flag bit corresponding to the head pointer position is the same as the flag bit corresponding to the load instruction position, determining the range after the head pointer position and before the load instruction position in the storage queue as the search range;

[0215] When the flag bit corresponding to the head pointer position is different from the flag bit corresponding to the load instruction position, the storage queue is rewound; and a range after the head pointer position and before the load instruction position in the storage queue after the rewinding is determined as the search range;

[0216] The flag bit is 0 or 1, and the flag bit switches when it moves from the end of the team to the head of the team.

[0217] In some embodiments, the dependency identifier includes a comparator array and a replicator, the number of comparators in the comparator array being the same as the number of entries in the storage queue;

[0218] The step of determining the instruction dependency bitmap by the dependency identifier based on the instruction address of the data storage instruction and the instruction address of the data loading instruction in the storage queue comprises:

[0219] The instruction address of the data storage instruction is compared with the instruction address of each of the data loading instructions one by one by the comparator array to obtain an original bitmap, wherein the original bitmap is composed of 0 and 1, wherein 0 is used to indicate that the instruction address of the data storage instruction is inconsistent with the instruction address of the data loading instruction, and 1 is used to indicate that the instruction address of the data storage instruction is consistent with the instruction address of the data loading instruction;

[0220] The original bitmap is copied by the replicator to obtain the instruction dependency bitmap, and the length of the instruction dependency bitmap is twice that of the original bitmap.

[0221] In some embodiments, each entry in the storage queue is n bytes, the forward data selection unit is composed of n forward data selection subunits, and different forward data selection subunits correspond to different bytes in the entry;

[0222] The n forward data selection subunits operate in parallel, and the first forward data is obtained by merging data output by each of the forward data selection subunits.

[0223] In some embodiments, for the i-th forward data selection subunit among the n forward data selection subunits, the i-th forward data selection subunit includes a second AND operator and a tree selector;

[0224] The selecting the first forward data from the storage queue by the forward data selection unit based on the dependent address sequence includes:

[0225] Performing a bitwise AND operation on the dependent address sequence and the i-th byte valid sequence by the second AND operator to obtain an i-th byte valid dependent address sequence, wherein the i-th byte valid sequence is used to indicate the validity of the data stored in the i-th byte in each entry;

[0226] Based on the i-th byte effective dependent address sequence, the i-th forward byte is selected from the i-th bytes of each entry in the storage queue by the tree selector.

[0227] In some embodiments, based on the i-th byte effective dependent address sequence, selecting the i-th forward byte from the i-th byte of each entry in the storage queue by the tree selector comprises:

[0228] Determining a valid i-th byte in each entry in the storage queue based on the i-th byte valid dependent address sequence;

[0229] Based on the merging rule, the valid i-th bytes in the adjacent entries are merged through the first-level tree structure, and the merging result is input into the second-level tree structure;

[0230] Based on the merging rule, adjacent merging results outputted from the j-1th layer of tree structure are merged through the jth layer of tree structure, and the merging results are inputted into the j+1th layer of tree structure, where j is an integer greater than or equal to 2.

[0231] In some embodiments, the merging rule indicates that a merging priority of the i-th byte in the k-th entry is higher than a merging priority of the i-th byte in the k-1-th entry.

[0232] In some embodiments, the i-th forward data selection subunit further includes a first replicator and a second replicator;

[0233] The method further comprises:

[0234] Copying the i-th byte valid sequence by the first replicator, wherein the length of the i-th byte valid sequence after copying is consistent with the length of the dependent address sequence;

[0235] The storage queue is replicated by the second replicator, wherein the length of the storage queue after replication is consistent with the length of the dependent address sequence.

[0236] The detailed process of data forwarding performed by the data forwarding device in the above-mentioned data forwarding method can be referred to the above-mentioned device embodiment, and this embodiment will not be described in detail here.

[0237] In some embodiments, the data forwarding device in the embodiments of the present application may be integrated in a chip. The embodiments of the present application provide a chip, the chip comprising the data forwarding device provided in any of the above embodiments.

[0238] Optionally, the chip may be a processor, such as an AI processor, a CPU processor, or other processor with data memory access requirements, which is not limited in the embodiments of the present application.

[0239] Please refer to Fig.11 , which shows a block diagram of a computer device 1200 provided by an exemplary embodiment of the present application. The computer device 1200 may be a portable mobile terminal, such as a smart phone, a tablet computer, a Moving Picture Experts Group Audio Layer III (MP3) player, or a Moving Picture Experts Group Audio Layer IV (MP4) player. The computer device 1200 may also be called a user device, a portable terminal, a workstation, a server, or other names.

[0240] Typically, the computer device 1200 includes a processor 1201 and a memory 1202 .

[0241] The processor 1201 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 1201 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), and programmable logic array (PLA). The processor 1201 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a central processing unit (CPU); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1201 may be integrated with a graphics processing unit (GPU), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1201 may also include an artificial intelligence (AI) processor, which is used to process computing operations related to machine learning.

[0242] In some embodiments, the processor 1201 may be integrated with the data forwarding device provided in the above embodiments. When there is a demand for data loading, the forwarding data may be obtained from the storage queue, storage buffer or data cache through the data forwarding device.

[0243] The memory 1202 may include one or more computer-readable storage media, which may be tangible and non-transitory. The memory 1202 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices, flash memory storage devices.

[0244] In some embodiments, the computer device 1200 may also optionally include a peripheral device interface 1203 and at least one peripheral device.

[0245] Those skilled in the art will understand that Fig.11 The structure shown in the figure does not constitute a limitation on the computer device 1200, and the computer device 1200 may include more or less components than those shown in the figure, or combine some components, or adopt a different arrangement of components.

[0246] A person skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware or by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.

[0247] The above description is only an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A data forwarding device, characterized in that: The device comprises: A storage queue, a storage buffer, a data cache, a first forward merging unit, and a second forward merging unit; The storage queue and the storage buffer are respectively connected to the first forward merging unit, and the second forward merging unit is respectively connected to the data cache and the first forward merging unit; The first forward merging unit is used to read the first forward data from the storage queue, read the second forward data from the storage buffer, and perform forward merging on the first forward data and the second forward data to obtain third forward data; The second forward merging unit is used to read cached data from the data cache based on the data load instruction, and read the third forward data from the first forward merging unit, and forward merge the third forward data and the cached data to obtain fourth forward data; The second forward merge unit is further configured to output the fourth forward data corresponding to the data loading instruction to the data loading pipeline.

2. The device according to claim 1, characterized in that The first forward merging unit is configured to select the data corresponding to the first address in the first forward data as the third forward data when the first forward data contains data corresponding to the first address and the second forward data does not contain data corresponding to the first address; and to select the data corresponding to the second address in the second forward data as the third forward data when the first forward data does not contain data corresponding to the second address and the second forward data contains data corresponding to the second address; The second forward merging unit is used to select the data corresponding to the third address in the third forward data as the fourth forward data when the third forward data contains data corresponding to the third address and the cache data does not contain data corresponding to the third address; and to select the data corresponding to the fourth address in the cache data as the fourth forward data when the third forward data does not contain data corresponding to the fourth address and the cache data contains data corresponding to the fourth address.

3. The device according to claim 2, characterized in that The forward delivery priority of the data in the storage queue is higher than the forward delivery priority of the data in the storage buffer, and the forward delivery priority of the data in the storage queue and the storage buffer is higher than the forward delivery priority of the data in the data cache; The first forward merging unit is further configured to select data belonging to the first forward data as the third forward data based on the overlapping addresses when there is an address overlap between the first forward data and the second forward data; The second forward merging unit is further configured to select data belonging to the third forward data as the fourth forward data based on the overlapping addresses when there is address overlap between the third forward data and the cache data.

4. The device according to claim 1, characterized in that The storage queue is provided with a dependent address determination unit and a forward data selection unit; The dependent address determining unit is used to determine a dependent address sequence based on the data loading instruction and the data storage instructions in the storage queue, wherein the dependent address sequence is used to indicate the instruction addresses of the data storage instructions on which the data loading instruction depends; The forward data selection unit is used to select the first forward data from the storage queue based on the dependent address sequence.

5. The device according to claim 4, characterized in that The dependency address determination unit includes a search range generator, a dependency relationship identifier, and a first AND operator; The search range generator is used to determine the search range in the storage queue based on the head pointer position of the storage queue and the load instruction position of the data load instruction, wherein the load instruction position is used to represent the instruction position of the data load instruction in the data storage instruction sequence flow; The dependency identifier is used to determine an instruction dependency bitmap based on the instruction address of the data storage instruction and the instruction address of the data loading instruction in the storage queue, wherein the instruction dependency bitmap is used to characterize the dependency relationship between the data loading instruction and each of the data storage instructions; The first AND operator is used to perform a bitwise AND operation on the search range and the instruction dependency bitmap to obtain the dependency address sequence.

6. The device according to claim 5, characterized in that The storage queue is a circular queue; The search range generator is used to determine the range after the head pointer position and before the load instruction position in the storage queue as the search range when the flag bit corresponding to the head pointer position is the same as the flag bit corresponding to the load instruction position; The search range generator is further configured to wrap the storage queue when the flag bit corresponding to the head pointer position is different from the flag bit corresponding to the load instruction position; Determine the range after the head pointer position and before the load instruction position in the storage queue after wrapping as the search range; The flag bit is 0 or 1, and the flag bit switches when it moves from the end of the team to the head of the team.

7. The device according to claim 5, characterized in that The dependency identifier includes a comparator array and a replicator, wherein the number of comparators in the comparator array is the same as the number of entries in the storage queue; The comparator array is used to compare the instruction address of the data storage instruction with the instruction address of each of the data loading instructions one by one to obtain an original bitmap, wherein the original bitmap is composed of 0 and 1, wherein 0 is used to indicate that the instruction address of the data storage instruction is inconsistent with the instruction address of the data loading instruction, and 1 is used to indicate that the instruction address of the data storage instruction is consistent with the instruction address of the data loading instruction; The replicator is used to replicate the original bitmap to obtain the instruction dependency bitmap, and the length of the instruction dependency bitmap is twice that of the original bitmap.

8. The device according to claim 4, characterized in that Each entry in the storage queue is n bytes, and the forward data selection unit is composed of n forward data selection sub-units, and different forward data selection sub-units correspond to different bytes in the entry; The n forward data selection subunits operate in parallel, and the first forward data is obtained by merging data output by each of the forward data selection subunits.

9. The device according to claim 8, characterized in that For the i-th forward data selection subunit among the n forward data selection subunits, the i-th forward data selection subunit comprises a second AND operator and a tree selector; The second AND operator is used to perform a bitwise AND operation on the dependent address sequence and the i-th byte valid sequence to obtain an i-th byte valid dependent address sequence, wherein the i-th byte valid sequence is used to indicate the validity of the data stored in the i-th byte in each entry; The tree selector is used to select the i-th forward byte from the i-th byte of each entry in the storage queue based on the i-th byte effective dependent address sequence.

10. The device according to claim 9, characterized in that The tree selector is used to: Determining a valid i-th byte in each entry in the storage queue based on the i-th byte valid dependent address sequence; Based on the merging rule, the valid i-th bytes in the adjacent entries are merged through the first-level tree structure, and the merging result is input into the second-level tree structure; Based on the merging rule, adjacent merging results outputted from the j-1th layer of tree structure are merged through the jth layer of tree structure, and the merging results are inputted into the j+1th layer of tree structure, where j is an integer greater than or equal to 2.

11. The device according to claim 10, characterized in that The merge rule indicates that the merge priority of the i-th byte in the k-th entry is higher than the merge priority of the i-th byte in the k-1-th entry.

12. The device according to claim 9, characterized in that The i-th forward data selection subunit further includes a first replicator and a second replicator; The first replicator is used to replicate the i-th byte valid sequence, wherein the length of the i-th byte valid sequence after replication is consistent with the length of the dependent address sequence; The second replicator is used to replicate the storage queue, wherein the length of the storage queue after replication is consistent with the length of the dependent address sequence.

13. A data forwarding method, characterized in that: The method comprises: Reading the first forward data from the storage queue and the second forward data from the storage buffer by the first forward merging unit, and forward merging the first forward data and the second forward data to obtain the third forward data; Based on the data load instruction, read the cache data from the data cache through the second forward merging unit, read the third forward data from the first forward merging unit, and forward merge the third forward data and the cache data to obtain fourth forward data; The fourth forward data corresponding to the data loading instruction is output to the data loading pipeline through the second forward merging unit.

14. A chip, characterized in that: The chip includes the data forwarding device as described in any one of claims 1 to 12.

15. A computer device, characterized in that: The computer device comprises a processor and a memory, wherein the processor is connected to the memory via a bus, and the processor is provided with a data forwarding device as claimed in any one of claims 1 to 12.

Citation Information

Cited By

  • Memory access unit, memory access instruction execution method and chip

    CN122240187A