Cross-architecture instruction processing method and processor based on pattern matching and fusion

Through a cross-architecture instruction processing method based on pattern matching and fusion, X86 instructions are converted into RISC-V instructions, which solves the problems of high resource consumption and low execution efficiency in the prior art, and realizes efficient and low resource consumption instruction processing, which is suitable for embedded systems and high real-time applications.

CN119718424BActive Publication Date: 2025-06-20SHANGHAI XINLIJI SEMICON CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510198917.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-20
Estimated Expiration
2045-02-24

AI Technical Summary

Technical Problem

When running X86 instructions on RISC architecture processors, the prior art requires a large amount of memory and computing resources, and the execution efficiency is low, which cannot meet the application scenarios with high real-time requirements.

Method used

A cross-architecture instruction processing method based on pattern matching and fusion is adopted to convert the instruction sequence in the first instruction set into the instruction sequence in the second instruction set. By matching the target first instruction sequence in the target instruction set, the corresponding target second instruction sequence is used instead, and if the match fails, the translation is performed.

Benefits of technology

It improves the processing efficiency of cross-architecture instructions, reduces resource consumption, reduces pipeline pauses, and improves the real-timeness of instruction processing. It is suitable for embedded systems with limited resources and application scenarios with high real-time requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119718424B_ABST
    Figure CN119718424B_ABST
Patent Text Reader

Abstract

The present invention discloses a cross-architecture instruction processing method and a processor based on pattern matching and fusion. The method is used to convert a first instruction sequence in a first instruction set into a second instruction sequence in a second instruction set. The method includes determining a target instruction set from the first instruction set, where the target instruction set includes multiple target first instruction sequences, determining target second instruction sequences used to replace the target first instruction sequences from the second instruction set, and the number of instructions included in the target second instruction sequences is not greater than the number of instructions included in the target first instruction sequences; obtaining the first instruction sequence to be processed, determining whether there is a target first instruction sequence in the target instruction set that matches the first instruction sequence, if so, replacing the first instruction sequence with the corresponding target second instruction sequence; if not, translating the first instruction sequence into a second instruction sequence. The present invention can improve the processing efficiency of cross-architecture instructions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technologies, and in particular, to a cross-architecture instruction processing method and a processor based on pattern matching and fusion. Background Art

[0002] In the field of computer science, the development of processor architectures has always been a hot topic. With the development of technologies, processors have evolved from a single CISC (Complex Instruction Set Computer) architecture to a RISC (Reduced Instruction Set Computer) architecture. The RISC architecture has received increasing attention due to its advantages such as simplicity, high efficiency, and modularity. However, in practical applications, due to historical reasons and the complexity of software ecosystems, many existing software and operating systems are designed based on the CISC architecture, such as the X86 architecture. Therefore, how to run X86 instructions on a RISC architecture processor has become an important issue.

[0003] In existing solutions, a widely used solution is to use Dynamic Binary Translation (DBT). The DBT technology runs an X86 emulator on a RISC architecture processor to translate X86 instructions into RISC-V instructions. This method can achieve the running of X86 instructions on a RISC architecture processor without modifying the existing software. However, this method requires a large amount of memory and computing resources, and due to the existence of the emulator, the execution efficiency is low.

[0004] Although the DBT technology can solve the problem of running X86 instructions on a RISC architecture processor to a certain extent, it has some significant limitations. First, the DBT technology requires a large amount of memory and computing resources, which is a huge challenge for resource-constrained embedded systems. Second, due to the existence of the emulator, the execution efficiency is low and cannot meet the application scenarios with high real-time requirements. In addition, the DBT technology may encounter performance bottlenecks when processing complex instruction sequences, affecting the overall performance of the system.

[0005] The disclosure of the above background art content is only used to assist in understanding the inventive concept and technical solution of the present invention, and it does not necessarily belong to the prior art of this application, nor will it necessarily give technical teachings; without clear evidence indicating that the above content was publicly available before the filing date of this application, the above background art should not be used to evaluate the novelty and inventiveness of this application. Summary of the Invention

[0006] The objective of the present invention is to provide a cross-architecture instruction processing method and a processor based on pattern matching and fusion, which can improve the processing efficiency of cross-architecture instructions and reduce resource consumption.

[0007] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0008] A cross-architecture instruction processing method based on pattern matching and fusion, which is used to convert a first instruction sequence in a first instruction set into a second instruction sequence in a second instruction set. The first instruction set and the second instruction set are based on different instruction set architectures. The first instruction sequence includes one or more instructions, and the second instruction sequence includes one or more instructions. The method includes the following steps:

[0009] Determine a target instruction set from the first instruction set. The target instruction set includes a plurality of target first instruction sequences, and the target first instruction sequences are part of the plurality of first instruction sequences. Determine a target second instruction sequence from the second instruction set for replacing the target first instruction sequence. The target second instruction sequence corresponds to the target first instruction sequence one by one, and the number of instructions included in the target second instruction sequence is not greater than the number of instructions included in the corresponding target first instruction sequence;

[0010] Obtain the first instruction sequence to be processed, and determine whether there is a target first instruction sequence in the target instruction set that matches the first instruction sequence. If there is, replace the first instruction sequence with the target second instruction sequence corresponding to the target first instruction sequence;

[0011] If not, translate the first instruction sequence into a second instruction sequence.

[0012] Further, based on any one of the above technical solutions or a combination of multiple technical solutions, the following steps are further included:

[0013] Classify the multiple target first instruction sequences in the target instruction set according to the function type to construct a function pattern library. The function pattern library includes multiple sub-pattern libraries, and the sub-pattern libraries include target first instruction sequences with the same function type. The target first instruction sequences included in different sub-pattern libraries correspond to different function types;

[0014] Configure a unique function type for each sub-pattern library, and the function type corresponding to the sub-pattern library corresponds to the function type corresponding to the target first instruction sequence included therein;

[0015] Judge whether there is a target first instruction sequence in the target instruction set that matches the first instruction sequence in the following manner:

[0016] Determine the function type corresponding to the first instruction sequence and use it as the first function type;

[0017] Determine whether there is a function type corresponding to the sub-pattern library that corresponds to the first function type. If so, match the first instruction sequence with the target first instruction sequence included in the sub-pattern library of the corresponding function type. If the match is successful, there is a target first instruction sequence in the target instruction set that matches the first instruction sequence;

[0018] If the match is unsuccessful or the function types corresponding to each sub-pattern library are not the first function type, there is no target first instruction sequence in the target instruction set that matches the first instruction sequence.

[0019] Further, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, the sub-pattern library includes one or more of a load-operate-store sub-pattern library, an inter-register operation sub-pattern library, an address calculation and memory access sub-pattern library, and a loop control sub-pattern library;

[0020] Among them, the load-operate-store sub-pattern library includes the following target first instruction sequences: a data loading instruction sequence from memory, an ALU execution instruction sequence, and a memory storing instruction sequence;

[0021] The inter-register operation sub-pattern library includes the following target first instruction sequences: a continuous register addition and subtraction operation instruction sequence, a shift instruction sequence, and a logical operation instruction sequence;

[0022] The memory access sub-pattern library includes the following target first instruction sequences: an address calculation instruction sequence based on an offset and a base address;

[0023] The loop control sub-pattern library includes the following target first instruction sequences: a conditional jump operation instruction sequence and a loop instruction sequence.

[0024] Further, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, it further includes obtaining the first instruction sequence to be processed in the following manner:

[0025] Determine the maximum length of the target first instruction sequence in the target instruction set;

[0026] Use a sliding window to detect the input first instruction stream in real time. The length of the sliding window is not less than the maximum length, and the first instruction stream includes multiple first instruction sequences.

[0027] Further, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, the adjacent two detection ranges of the sliding window are adjacent or have an overlapping part.

[0028] Further, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, the first instruction sequence is translated into a second instruction sequence in the following manner:

[0029] A parallel decoding unit is used to perform parallel processing and field separation on multiple instructions in the first instruction sequence, including extracting an operation code, operands, and prefix information;

[0030] Each instruction in the first instruction sequence is translated through a look-up table or a preset translation rule, and the second instruction sequence is determined according to the translation results of the respective instructions.

[0031] Further, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, translating the first instruction sequence into a second instruction sequence further includes:

[0032] Configure a cache unit, which is configured to store multiple prior first instruction sequences and prior second instruction sequences for replacing the prior first instruction sequences. The first instruction set includes the prior first instruction sequences, and the second instruction set includes the prior second instruction sequences;

[0033] The usage frequency of the prior first instruction sequence is higher than a preset first usage frequency, and / or the time period from the most recent usage time of the prior first instruction sequence to the current time is within a preset time period range;

[0034] Determine whether the first instruction sequence is one of the multiple prior first instruction sequences. If so, directly replace the first instruction sequence with the corresponding prior second instruction sequence;

[0035] A preferred method is that determining whether the first instruction sequence is one of the multiple prior first instruction sequences and translating each instruction through a look-up table or a preset translation rule in parallel can effectively shorten the instruction translation time and improve the instruction translation efficiency;

[0036] Another preferred method is that determining whether the first instruction sequence is one of the multiple prior first instruction sequences before translating the first instruction sequence into a second instruction sequence can improve the translation efficiency while greatly reducing the resource consumption.

[0037] Further, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, translating the first instruction sequence into a second instruction sequence further includes:

[0038] Identify the first instruction sequence that does not conform to the specification through the legality check of the operation code and operands and issue a prompt message.

[0039] Further, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, the method is used to convert a first instruction stream into a second instruction stream. The first instruction stream includes a plurality of the first instruction sequences, and the second instruction stream includes a plurality of the second instruction sequences.

[0040] Further, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, for the current first instruction sequence in the first instruction stream, if there is no target first instruction sequence matching the current first instruction sequence in the target instruction set, the current first instruction sequence is translated into a second instruction sequence;

[0041] While translating the current first instruction sequence, for the first instruction sequence in the first instruction stream after the current first instruction sequence, it is determined whether there is a target first instruction sequence matching the first instruction sequence in the target instruction set.

[0042] Further, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, the first instruction set is based on the X86 architecture, and the second instruction set is based on the RISC-V architecture; or,

[0043] The first instruction set is based on the ARM architecture, and the second instruction set is based on the RISC-V architecture.

[0044] Further, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, the usage frequency of the target first instruction sequence is higher than a preset second usage frequency. That is, preferably, the first instruction sequences with high usage frequencies can be screened out according to the usage frequencies of the first instruction sequences in the first instruction set, and then it is further determined whether the first instruction sequences with high usage frequencies can be fused into second instruction sequences with fewer instructions. If so, the first instruction sequence is determined as the target first instruction sequence, and the set of all obtained target first instruction sequences is determined as the target instruction set.

[0045] According to another aspect of the present invention, the present invention provides a cross-architecture instruction processor based on pattern matching and fusion, which is used to convert instruction sequences in a first instruction set into instruction sequences in a second instruction set. The first instruction set and the second instruction set are based on different instruction set architectures;

[0046] The cross-architecture instruction processor based on pattern matching and fusion includes an instruction fusion hardware module and a translation module; the instruction fusion hardware module is configured to determine whether there is a target first instruction sequence matching the first instruction sequence to be processed in the target instruction set. If so, the target second instruction sequence corresponding to the target first instruction sequence is used to replace the first instruction sequence;

[0047] If not, the translation module is utilized to translate the first instruction sequence into a second instruction sequence, and the second instruction sequence includes one or more instructions in the second instruction set;

[0048] Wherein, the target instruction set includes a plurality of the target first instruction sequences, the target first instruction sequence includes one or more instructions in the first instruction set, and the target first instruction sequence satisfies the following conditions:

[0049] There exists a target second instruction sequence in the second instruction set for replacing the target first instruction sequence, and the number of instructions included in the target second instruction sequence is not greater than the number of instructions included in the corresponding target first instruction sequence.

[0050] Further, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, a function mode library is configured in the instruction fusion hardware module, the function mode library includes a plurality of sub-mode libraries, the sub-mode libraries include the target first instruction sequences with the same function type, and the function types corresponding to the target first instruction sequences included in different sub-mode libraries are different;

[0051] Each sub-mode library is configured with a unique function type, and the function type corresponding to the sub-mode library corresponds to the function type corresponding to the target first instruction sequence included therein;

[0052] The instruction fusion hardware module is further configured to determine whether there exists a target first instruction sequence in the target instruction set that matches the first instruction sequence in the following manner:

[0053] Determine the function type corresponding to the first instruction sequence and use it as the first function type;

[0054] Judge whether there exists a function type corresponding to the sub-mode library that corresponds to the first function type. If so, match the first instruction sequence with the target first instruction sequences included in the sub-mode library corresponding to the function type. If the match is successful, there exists a target first instruction sequence in the target instruction set that matches the first instruction sequence;

[0055] If the match is unsuccessful or the function types corresponding to each sub-mode library are not the first function type, there does not exist a target first instruction sequence in the target instruction set that matches the first instruction sequence.

[0056] Further, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, the translation module includes a hardware translator and a cache unit, and the hardware translator is configured to translate the first instruction sequence into a second instruction sequence;

[0057] The cache unit is configured to store a plurality of prior first instruction sequences and prior second instruction sequences for replacing the prior first instruction sequences. The first instruction set includes the prior first instruction sequences, and the second instruction set includes the prior second instruction sequences;

[0058] The usage frequency of the prior first instruction sequence is higher than a preset first usage frequency, and / or the time period from the last usage moment of the prior first instruction sequence to the current moment is within a preset time period range;

[0059] Use the hardware translator to translate the first instruction sequence into a second instruction sequence, and determine whether the first instruction sequence is one of the plurality of prior first instruction sequences. If so, directly replace the first instruction sequence with the corresponding prior second instruction sequence; if not, use the translation result obtained by translating the first instruction sequence by the hardware translator as the second instruction sequence.

[0060] The beneficial effects brought by the technical solution provided by the present invention are as follows:

[0061] a. The present invention realizes the processing of the first instruction sequence to the second instruction sequence through the cooperation of instruction matching and fusion and instruction translation, reduces the use of the simulator, thereby improving the efficiency of instruction execution. Compared with the traditional DBT technology, this solution can run X86 instructions on a RISC architecture processor under more efficient conditions, and applying this method to the pipeline of the processor for cross-architecture instruction processing also does not require additional memory and computing resources, can reduce pipeline stalls, improve the real-time performance of instruction processing, has significant advantages in resource consumption, is more suitable for application scenarios with high real-time requirements compared with DBT technology, and is also particularly suitable for embedded systems with limited resources;

[0062] b. The present invention optimizes the pattern matching and fusion of the first instruction sequence or the first instruction stream, can directly convert the corresponding first instruction sequence into a second instruction sequence with fewer instructions, avoids the performance bottleneck of the cross-architecture instruction processing system, improves the overall performance of the system, and can also perform better when processing complex instruction sequences compared with the traditional DBT technology;

[0063] c. The present invention sets a cache unit to store the recently used first instruction sequence and / or the frequently used first instruction sequence. If the cache unit hits the first instruction sequence to be translated currently, the previously translated second instruction sequence can be directly output, which can reduce the overhead caused by repeated translation, improve the translation efficiency of the first instruction and reduce the consumption of resources. Description of the Drawings

[0064] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments described in the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0065] Figure 1 Flowchart of a cross-architecture instruction processing method based on pattern matching and fusion provided for an exemplary embodiment of the present invention;

[0066] Figure 2 Schematic diagram of the process of processing an X86 instruction stream into a RISC-V instruction stream provided for an exemplary embodiment of the present invention;

[0067] Figure 3 Flowchart of the first method for translating a first instruction sequence provided for an exemplary embodiment of the present invention;

[0068] Figure 4 Flowchart of the second method for translating a first instruction sequence provided for an exemplary embodiment of the present invention;

[0069] Figure 5 Block diagram of a cross-architecture instruction processor based on pattern matching and fusion provided for an exemplary embodiment of the present invention. Detailed implementation manners

[0070] In order to enable those skilled in the art of the present technology to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0071] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or equipment comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or equipment.

[0072] Based on the deficiencies of the prior art, the present application aims to provide a cross-architecture instruction processing method and a processor based on pattern matching and fusion, so as to achieve cross-architecture instruction processing and operation with high efficiency and low resource consumption, including running the X86 instruction set on the RISC-V architecture, running the ARM instruction set on the RISC-V architecture, etc., which can promote the development of the instruction set architecture towards high performance, high efficiency, low power consumption and multi-architecture fusion.

[0073] In an embodiment of the present invention, a cross-architecture instruction processing method based on pattern matching and fusion is provided, which is used to convert a first instruction sequence in a first instruction set into a second instruction sequence in a second instruction set. The first instruction set and the second instruction set are based on different instruction set architectures. The first instruction sequence includes one or more instructions, and the second instruction sequence includes one or more instructions. Among them, the first instruction set is based on the X86 architecture, and the second instruction set is based on the RISC-V architecture to achieve cross-architecture operation of the X86 instruction set based on the RISC-V architecture; it can also be that the first instruction set is based on the ARM architecture, and the second instruction set is based on the RISC-V architecture to achieve cross-architecture operation of the ARM instruction set based on the RISC-V architecture. It should be noted that this method is not limited to these two cross-architecture instruction processing methods, and is also applicable to other cross-architecture instruction processing situations.

[0074] In this embodiment, referring to Figure 1 , the cross-architecture instruction processing method based on pattern matching and fusion includes the following steps:

[0075] Determine a target instruction set from the first instruction set. The target instruction set includes multiple target first instruction sequences. Determine a target second instruction sequence from the second instruction set for replacing the target first instruction sequence. The number of instructions included in the target second instruction sequence is not greater than the number of instructions included in the target first instruction sequence, so as to fuse / convert a first instruction sequence with a larger number of instructions into a second instruction sequence with a smaller number of instructions (at least not more). For example, fuse (combine) three X86 instructions (LOAD, ALU operation, and STORE) into one composite RISC-V instruction;

[0076] Obtain the first instruction sequence to be processed, and determine whether there is a target first instruction sequence in the target instruction set that matches the first instruction sequence. If so, replace the first instruction sequence with the target second instruction sequence corresponding to the target first instruction sequence;

[0077] If not, translate the first instruction sequence into a second instruction sequence.

[0078] Among them, through a pre-determined target instruction set and pre-determined target second instruction sequences used to replace each target first instruction sequence in the target instruction set, through matching judgment, directly replacing the successfully matched first instruction sequence with the corresponding target second instruction sequence and outputting it can effectively improve the efficiency and accuracy of cross-architecture instruction processing / translation.

[0079] In addition, since the number of instructions included in the target second instruction sequence is not greater than the number of instructions included in the target first instruction sequence, therefore, through this method, it is possible to match and fuse the first instruction sequence into a second instruction sequence with a smaller number. Compared with the existing dynamic binary translation (DBT) technology, it can effectively reduce the number of translated instructions. Compared with the existing technology, this method has significant advantages in improving the execution efficiency of cross-architecture instructions, reducing resource consumption, optimizing the instruction stream, improving real-time performance and compatibility, etc.

[0080] In an embodiment of the present invention, in order to further improve the execution efficiency of cross-architecture instruction processing, reduce resource consumption, and avoid performance bottlenecks, a cross-architecture instruction processing method based on pattern matching and fusion is further implemented through the following method.

[0081] See Figure 1 and Figure 2 , classify multiple target first instruction sequences in the target instruction set according to the function type to construct a function pattern library. The function pattern library includes multiple sub-pattern libraries, and the sub-pattern libraries include the target first instruction sequences with the same function type. The target first instruction sequences included in different sub-pattern libraries correspond to different function types.

[0082] Configure a unique function type for each sub-pattern library, and the function type corresponding to the sub-pattern library corresponds to the function type corresponding to the target first instruction sequence it includes. It should be noted that the corresponding here includes the same, consistent, and matching.

[0083] Taking the X86 instruction set or the ARM instruction set as an example, the function types of the first instruction sequences it includes usually include load-operation-store (LOAD-OP-STORE), inter-register operation, address calculation and memory access, and loop control, etc. Among them, loading usually refers to loading data from memory to a register (for example, MOV EAX, [addr]), operation usually refers to operating on the loaded data (such as addition, subtraction, etc., for example, ADD EAX, EBX), and storing usually refers to storing the processed data back to memory (for example, MOV [addr], EAX).

[0084] Correspondingly, the function mode library includes, but is not limited to, one or more of the following sub-mode libraries: load-operation-store sub-mode library, inter-register operation sub-mode library, address calculation and memory access sub-mode library, and loop control sub-mode library.

[0085] Among them, the load-operation-store sub-mode library includes the following target first instruction sequences: data loading instruction sequence from memory, ALU (Arithmetic and Logic Unit) instruction execution sequence, and data storing-back instruction sequence to memory. Correspondingly, the function type of the load-operation-store sub-mode library is determined to be load-operation-store.

[0086] The inter-register operation sub-mode library includes the following target first instruction sequences: consecutive register addition and subtraction operation instruction sequence, shift instruction sequence, and logical operation instruction sequence. Correspondingly, the function type of the inter-register operation sub-mode library is determined to be inter-register operation.

[0087] The memory access sub-mode library includes the following target first instruction sequences: address calculation instruction sequence based on offset and base address. Correspondingly, the function type of the memory access sub-mode library is determined to be memory access.

[0088] The loop control sub-mode library includes the following target first instruction sequences: conditional jump operation instruction sequence and loop instruction sequence. Correspondingly, the function type of the loop control sub-mode library is determined to be loop control.

[0089] In this embodiment, it is not necessary to compare the first instruction sequence to be processed one by one with each target first instruction sequence in the target instruction set to determine whether there is a target first instruction sequence that matches the first instruction sequence. Instead, efficient pattern matching is performed through the following method:

[0090] Determine the function type corresponding to the first instruction sequence and use it as the first function type;

[0091] Judge whether there is a function type corresponding to the sub-mode library that corresponds to the first function type. If so, match the first instruction sequence with the target first instruction sequences included in the sub-mode library. If the match is successful, there is a target first instruction sequence in the target instruction set that matches the first instruction sequence;

[0092] If the match is unsuccessful or the function types corresponding to each sub-mode library are not the first function type, there is no target first instruction sequence in the target instruction set that matches the first instruction sequence.

[0093] It should be noted that the names and functional types of the respective sub-pattern libraries in this embodiment are only for illustrative purposes, and the protection scope of the present application is not limited by the names and functional types of the respective sub-pattern libraries. In other embodiments, the names of the respective sub-pattern libraries and their corresponding functional types may also be determined by means such as numbers, labels, etc. For example, sub-pattern library 1, sub-pattern library 2,..., sub-pattern library n (n is an integer not less than 2), and the corresponding functional types are determined as type 1, type 2,..., type n. As long as the multiple target first instruction sequences in the target instruction set are classified according to the functional type to construct a functional pattern library including multiple sub-pattern libraries, and the target first instruction sequences with the same functional type are included in the same sub-pattern library, and the functional types corresponding to the target first instruction sequences included in different sub-pattern libraries are different, they all fall within the protection scope of the present application.

[0094] The present application proposes to classify the multiple target first instruction sequences in the target instruction set according to the functional type, rather than other methods such as usage frequency, because based on the functional type classification, the functional type corresponding to the first instruction sequence can be quickly determined by extracting the operation code of the first instruction sequence to be processed, thereby narrowing the matching range of the first instruction sequence. If classified according to the instruction length or usage frequency, etc., the functional pattern library cannot be matched by the operation code of the first instruction sequence. However, when determining the target instruction set from the first instruction set, the first instruction sequences with high frequency usage (usage frequency higher than the preset second usage frequency) can be screened out according to the usage frequency, and among the first instruction sequences with high frequency usage, the first instruction sequences that can be merged into second instruction sequences with fewer (or equivalent) instruction numbers can be further screened out, thereby obtaining the target instruction set.

[0095] The method is not only applicable to converting the first instruction sequence into the second instruction sequence, but also applicable to converting the first instruction stream into the second instruction stream. Wherein, the first instruction stream includes multiple first instruction sequences, and the second instruction stream includes multiple second instruction sequences.

[0096] In an embodiment of the present invention, the first instruction stream is timely detected, matched, and translated in the following manner to output the second instruction stream in real time. As Figure 3 shown, a sliding window detection mechanism is adopted to obtain the first instruction sequence to be processed. Through the sliding window, the input instruction stream can be analyzed in real time and it can be detected in real time whether the input first instruction stream, such as the X86 instruction stream, conforms to the pattern in the pattern library. The function of the sliding window is to maintain a queue of instruction streams of a fixed length and gradually slide to check whether there is a sequence matching the pattern library.

[0097] Preferably, the size of the sliding window can be dynamically adjusted, that is, the size (i.e., length) of the sliding window is defined according to the length of the longest first instruction sequence in the function mode library, and the detection ranges of two adjacent sliding windows are adjacent or have an overlapping part to ensure that all possible first instruction sequences are covered.

[0098] Therefore, before using the sliding window to perform matching detection on the first instruction stream, it is first necessary to determine the maximum length of the target first instruction sequence in the target instruction set. Then, the length of the sliding window is defined to be not less than the maximum length. The first instruction stream includes a plurality of the first instruction sequences, usually a plurality of first instruction sequences input sequentially.

[0099] Matching algorithm: Each time the sliding window moves, the instruction sequence within the window is matched and compared with each sub-mode library in the above function mode library. Inside the sliding window, according to the operation code of the extracted instruction sequence, the sub-mode library with the same function type is determined. Preferably, hardware parallel logic is used to simultaneously match multiple sub-mode libraries. If there is a sub-mode library with the same function type as the instruction sequence extracted by the sliding window, the instruction sequence is compared with the target first instruction sequence in the sub-mode library. If they are the same, the matching is successful. If the matching is successful, it directly enters the instruction fusion logic step, that is, directly replaces the first instruction sequence with the corresponding target second instruction sequence and outputs it. If the matching fails, the first instruction sequence is translated into a second instruction sequence by means of hardware translation or software translation (such as the above DBT technology).

[0100] It should be noted that in this embodiment, as Figure 1 and Figure 2 shown, for the current first instruction sequence in the first instruction stream, if there is no target first instruction sequence in the target instruction set that matches the current first instruction sequence, the current first instruction sequence is translated into a second instruction sequence. And when translating the current first instruction sequence, it is simultaneously determined whether there is a target first instruction sequence in the target instruction set that matches the first instruction sequence after the current first instruction sequence in the first instruction stream. That is, for the first instruction sequence that fails to match, it will be directly passed to the translation logic process without interfering with the normal operation of the pattern matching and instruction fusion logic. That is, the pattern matching judgment and instruction fusion output process of the first instruction sequence obtained by the sliding window subsequently can be synchronized with the translation process of the first instruction sequence that failed to match previously. Therefore, the efficiency of cross-architecture instruction processing, running, and execution can be improved more efficiently.

[0101] In an embodiment of the present invention, a method for accelerating the translation efficiency of the first instruction sequence based on hardware is provided, as Figure 3As shown, a cache unit is pre-configured. The cache unit is configured to store a plurality of prior first instruction sequences and prior second instruction sequences for replacing the prior first instruction sequences. The first instruction set includes the prior first instruction sequences, and the second instruction set includes the prior second instruction sequences. It should be noted that the prior first instruction sequences are the first instruction sequences that failed to match previously and entered the translation process, and these prior first instruction sequences also meet the following conditions: the usage frequency of the prior first instruction sequences is higher than a preset first usage frequency, and / or the time period from the most recent usage moment of the prior first instruction sequences to the current moment is within a preset time period range. The prior second instruction sequences are the translation results corresponding to the prior first instruction sequences. By setting the cache unit, the overhead caused by repeated translation can be reduced, the translation efficiency can be improved, and the consumption of resources can be reduced.

[0102] Preferably, the cache unit is set to a two-level cache structure, including a first-level cache structure and a second-level cache structure. Among them, the first-level cache structure stores the most recent first instruction sequences and their translation results; the second-level cache structure stores the frequently used first instruction sequences and their translation results.

[0103] Judge whether the first instruction sequence is one of the plurality of prior first instruction sequences in the two-level cache structure. If so, directly replace the first instruction sequence with the corresponding prior second instruction sequence; if not, translate the first instruction sequence into a second instruction sequence in the following manner:

[0104] Use a parallel decoding logic to perform parallel processing and field separation on multiple instructions in the first instruction sequence, including extracting operation codes, operands, and prefix information. Among them, the parallel decoding logic supports parsing multiple bytes simultaneously, and the decoding speed of the first instruction sequence is improved through a pipelined design.

[0105] Translate each of the instructions through a look-up table (LUT, Look-Up-Table) or a preset translation rule, and determine the second instruction sequence according to the translation results of each of the instructions.

[0106] Realize the mapping from X86 to RISC-V instructions through a look-up table (LUT). Its translation process includes: searching according to the rule of retrieving the corresponding RISC-V instruction template from the LUT according to the operation code of the first instruction sequence; for the processing of complex instructions, for example, for complex X86 instructions that require multiple RISC-V instructions to complete, design a multi-stage translation logic to gradually generate a complete second instruction sequence.

[0107] As Figure 3The instruction translation process shown first queries the cache unit. If there is no previous translation result for the current first instruction sequence, the current first instruction sequence is translated using a lookup table, which can balance translation efficiency and resource consumption. In another embodiment of the present invention, different from Figure 3 the process shown, in this embodiment, the translation process of the first instruction sequence is as shown in Figure 4 When translating the current first instruction sequence, the cache unit and the LUT are queried in parallel. If a previous translation record is found in the cache unit, the translation result is directly output and the lookup table translation method is terminated. If no previous translation record is found in the cache unit, the translation result is obtained using the lookup table. This process consumes more resources, but can greatly shorten the translation time and improve the translation efficiency of the first instruction sequence.

[0108] When translating the first instruction sequence into the second instruction sequence using the lookup table, it further includes: identifying the first instruction sequence that does not conform to the specification through the legality check of the operation code and the operand and sending a prompt message. Lookup table (LUT) match failure: If an instruction cannot find a matching item in the LUT, the instruction is considered illegal and an exception signal is triggered.

[0109] The following uses specific application examples to illustrate the improvement of the cross-architecture instruction translation efficiency and processing efficiency of the cross-architecture instruction processing method based on pattern matching and fusion proposed in this application.

[0110] For an X86 instruction sequence (instruction sequence 1) with a function type of the described load-operate-store, the first instruction sequence 1 is:

[0111] MOV EAX, [0×1000]; Load data from memory to EAX;

[0112] ADD EAX, EBX; EAX = EAX + EBX;

[0113] MOV [0×1000], EAX; Store the value of EAX back to memory;

[0114] The RISC-V instruction sequence translated (without fusion) from this X86 instruction sequence (instruction sequence 1) is as follows:

[0115] LW t0, 0×1000(t1); Load memory data to t0;

[0116] ADD t0, t0, t2; t0 = t0 + t2;

[0117] SW t0, 0×1000(t1); Store the value of t0 back to memory;

[0118] The X86 instruction sequence (Instruction Sequence 1) is fused to obtain the RISC-V instruction sequence (Target Second Instruction Sequence 1) as: AMOADD W t0, t0, (0x1000); perform an atomic addition operation on the address 0x1000.

[0119] It can be seen that by using the pattern matching and fusion method, the Target Second Instruction Sequence 1 with a shorter length and more concise form can be directly used to replace the First Instruction Sequence 1, which not only does not require a long translation time, but also can reduce the instruction length and improve the instruction execution efficiency.

[0120] For an X86 instruction sequence (Instruction Sequence 2) with a function type of address calculation and memory access, the First Instruction Sequence 2 is:

[0121] LEA EAX, [EBX + ECX * 4]; load the address of EBX + ECX * 4 into EAX;

[0122] MOV [EAX], EDX; store the value of EDX into the calculated address.

[0123] The RISC-V instruction sequence translated (without fusion) from this X86 instruction sequence (Instruction Sequence 2) is as follows:

[0124] ADD t0, tl, t2; t0 = tl + t2;

[0125] SLL t0, t0, 2; t0 = t0 * 4

[0126] SW t3, 0(t0); store the value of t3 into the address calculated by t0.

[0127] The X86 instruction sequence (Instruction Sequence 2) is fused to obtain the RISC-V instruction sequence (Target Second Instruction Sequence 2) as: AMOADD.W t0, t3, (t1 + t2 * 4); perform an addition operation at the address t1 + t2 * 4.

[0128] It can be seen that if the translation method is adopted, Instruction Sequence 2 is usually translated into a longer RISC-V instruction sequence, but by using the pattern matching and fusion method, the Target Second Instruction Sequence 2 with a shorter length can be directly used to replace the First Instruction Sequence 2, which not only does not require a long translation time, but also can reduce the instruction length and improve the instruction execution efficiency.

[0129] In an embodiment of the present invention, a cross-architecture instruction processor based on pattern matching and fusion is provided, which is used to convert an instruction sequence in a first instruction set into an instruction sequence in a second instruction set, and the first instruction set and the second instruction set are based on different instruction set architectures.

[0130] The cross-architecture instruction processor based on pattern matching and fusion includes an instruction fusion hardware module (hereinafter referred to as IFHM) and a translation module (hereinafter referred to as HTA). The instruction fusion hardware module and the translation module are configured to process the first instruction stream in parallel, and the priority of the instruction fusion hardware module is higher than that of the translation module.

[0131] As Figure 5 shown, the cross-architecture instruction processor based on pattern matching and fusion includes a decoder, an instruction fusion hardware module, and a translation module. The first instruction stream input to the processor, such as the X86 instruction stream, first passes through the decoder for decoding and field separation, and the decoded X86 instruction stream can be obtained and output to the instruction fusion hardware module.

[0132] The instruction fusion hardware module includes a sliding window detection module, a pattern matching module, and a functional pattern library. Among them, the functional pattern library stores a target instruction set, the target instruction set includes a plurality of the target first instruction sequences, the target first instruction sequence includes one or more instructions in the first instruction set, and the target first instruction sequence meets the following conditions:

[0133] There is a target second instruction sequence in the second instruction set for replacing the target first instruction sequence, and the number of instructions included in the target second instruction sequence is not greater than the number of instructions included in the corresponding target first instruction sequence.

[0134] More preferably, the multiple target first instruction sequences in the target instruction set are classified according to function types to construct corresponding sub-pattern libraries, a unique function type is configured for each sub-pattern library, and the function type corresponding to the sub-pattern library corresponds to the function type corresponding to the target first instruction sequence included therein. Therefore, the functional pattern library includes multiple sub-pattern libraries, the sub-pattern library includes multiple target first instruction sequences with the same function type, and the function types corresponding to the target first instruction sequences included in different sub-pattern libraries are different.

[0135] The sliding window detection module is configured to analyze the input instruction stream in real time using a sliding window, that is, to detect the input X86 instruction stream in real time. Its function is to maintain an instruction stream queue of a fixed length, gradually slide to extract the corresponding instruction sequence, and cooperate with the pattern matching module to perform pattern matching judgment on the current first instruction sequence, so as to quickly identify the first instruction sequence that can be fused in the first instruction stream.

[0136] The pattern matching module is configured to determine whether there is a target first instruction sequence in the target instruction set that matches the first instruction sequence to be processed. If so, the target first instruction sequence is used to replace the first instruction sequence with the corresponding target second instruction sequence. If not, the first instruction sequence for which pattern matching fails is transmitted to the translation module, and the translation module is used to translate the first instruction sequence into the corresponding second instruction sequence.

[0137] Based on a functional pattern library including multiple sub-pattern libraries, the instruction fusion hardware module is further configured to determine whether there is a target first instruction sequence in the target instruction set that matches the first instruction sequence in the following manner:

[0138] Determine the functional type corresponding to the first instruction sequence as the first functional type;

[0139] Determine whether there is a functional type corresponding to the sub-pattern library that corresponds to the first functional type. If so, the first instruction sequence is matched with the target first instruction sequence included in the sub-pattern library. If the match is successful, there is a target first instruction sequence in the target instruction set that matches the first instruction sequence.

[0140] If the match is unsuccessful or the functional types corresponding to each sub-pattern library do not correspond to the first functional type, there is no target first instruction sequence in the target instruction set that matches the first instruction sequence.

[0141] The instruction fusion hardware module stores common / high-frequency used X86 instruction sequences through a functional pattern library, and according to predefined fusion rules, fuses the matched X86 instruction sequences into efficient RISC-V composite instructions, that is, converts multiple instructions into efficient RSIC-V instructions, thereby reducing the number of instructions and optimizing the execution efficiency.

[0142] The translation module includes a hardware translator and a cache unit. The hardware translator is configured to translate the first instruction sequence into a second instruction sequence.

[0143] The cache unit stores previous translation results. The cache unit is configured to store multiple previous first instruction sequences and the previous second instruction sequences used to replace the previous first instruction sequences. The first instruction set includes the previous first instruction sequences, and the second instruction set includes the previous second instruction sequences.

[0144] Among them, the usage frequency of the prior first instruction sequence is higher than a preset first usage frequency, and / or the time period from the most recent usage time of the prior first instruction sequence to the current time is within a preset time period range. Preferably, the cache unit has a two-level cache structure, including a first-level cache structure and a second-level cache structure. Among them, the first-level cache structure stores the most recent first instruction sequence and its translation result; the second-level cache structure stores the frequently used first instruction sequence and its translation result.

[0145] Use the hardware translator to translate the first instruction sequence into a second instruction sequence, and determine whether the first instruction sequence is one of the multiple prior first instruction sequences. If so, directly use the corresponding prior second instruction sequence to replace the first instruction sequence; if not, use the translation result obtained by the hardware translator to translate the first instruction sequence as the second instruction sequence. It should be noted that it can be first determined whether there is a translation record in the cache unit that hits the current first instruction sequence. If so, directly call and output it; if not, then use the hardware translator to translate the first instruction sequence. It is also possible to set the search for the prior translation record using the cache unit and the translation using the hardware translator to be carried out synchronously to further improve the translation efficiency of the first instruction sequence.

[0146] Among them, the design of the cache unit can improve the translation performance of the translation module and reduce the overhead of repeated translation. The hardware translator uses a multi-byte parallel decoding technology to quickly parse X86 instructions, including information such as operation codes, operands, and prefixes, and uses a lookup table to complete instruction mapping, and translates it into an equivalent RISC-V instruction sequence.

[0147] The proposed cross-architecture instruction processor based on pattern matching and fusion described in the above embodiments adopts a hardware modular design method, and through the collaborative work of the instruction fusion hardware module and the translation module, realizes efficient instruction translation and execution optimization from X86 to RISC-V.

[0148] In summary, the present invention forms an efficient X86 to RISC-V translation and optimized execution architecture through technical means such as the collaborative working mechanism of hardware translation acceleration and instruction fusion optimization and modular pipeline integration. The combination of these technical means greatly improves the instruction execution efficiency and reduces the power consumption, providing strong hardware support for the migration of the X86 ecosystem to RISC-V.

[0149] Therefore, it can have a wide range of applications in the fields of software development, embedded systems, servers, and cloud computing. In terms of software development, the technical solutions described in the above embodiments can efficiently fuse and translate X86 instructions into RISC-V instructions, which can greatly improve the efficiency of software development. Moreover, developers can use existing X86 development tools and libraries without worrying about the differences in underlying hardware architectures. In addition, since the technical solutions of the present invention can achieve relatively high execution efficiency for cross-architecture instruction processing, it can meet application scenarios with relatively high real-time requirements, such as game development and multimedia processing. Applying the present technical solution to an embedded system can efficiently run X86 instructions in an embedded system with limited resources, which is a huge advantage for an embedded system with limited resources. For example, in applications such as smart homes and smart wearable devices, the present technical solution can provide high-performance and low-power computing support to meet the requirements of the devices for real-time performance and energy consumption. In the fields of servers and cloud computing, high performance and low latency are crucial. The present technical solution can significantly improve the performance of servers, reduce latency, and improve the overall efficiency of the system by optimizing the instruction stream and execution efficiency. In addition, since the execution efficiency of the present technical solution is relatively high, it can reduce the number of servers, lower the operating costs, and improve the economic benefits. Generally speaking, the present technical solution has broad application prospects in the fields of software development, embedded systems, servers, and cloud computing, meeting the market's demand for efficient, low-power, and high-performance computing.

[0150] It should be noted that the embodiments of the cross-architecture instruction processor based on pattern matching and fusion provided by the present invention have the same inventive concept as the embodiments of the cross-architecture instruction processing method based on pattern matching and fusion. The entire content of the embodiments of the cross-architecture instruction processing method based on pattern matching and fusion is incorporated into the embodiments of the cross-architecture instruction processor based on pattern matching and fusion by reference.

[0151] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device including the element.

[0152] The above are only specific embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A cross-architecture instruction processing method based on pattern matching and fusion, characterized in that: The method is used to convert a first instruction sequence in a first instruction set into a second instruction sequence in a second instruction set, wherein the first instruction set and the second instruction set are based on different instruction set architectures, and the method comprises the following steps: Determine a target instruction set from the first instruction set, the target instruction set includes a plurality of target first instruction sequences, classify the plurality of target first instruction sequences in the target instruction set according to function types to construct a function mode library, the function mode library includes a plurality of sub-mode libraries, the sub-mode libraries include the target first instruction sequences with the same function types, and the function types corresponding to the target first instruction sequences included in different sub-mode libraries are different; A unique function type is configured for each of the sub-pattern libraries, and the function type corresponding to the sub-pattern library corresponds to the function type corresponding to the target first instruction sequence included therein; Determine a target second instruction sequence from the second instruction set for replacing the target first instruction sequence, wherein the number of instructions included in the target second instruction sequence is not greater than the number of instructions included in the corresponding target first instruction sequence; Acquire a first instruction sequence to be processed, determine whether there is a target first instruction sequence matching the first instruction sequence in the target instruction set, and if so, replace the first instruction sequence with the target second instruction sequence corresponding to the target first instruction sequence; if not, translate the first instruction sequence into a second instruction sequence; It is determined whether the target first instruction sequence that matches the first instruction sequence exists in the target instruction set by: Determine a function type corresponding to the first instruction sequence and use it as a first function type; Determine whether there is a function type corresponding to the sub-pattern library corresponding to the first function type; if so, match the first instruction sequence with a target first instruction sequence included in the sub-pattern library of the corresponding function type; if the match is successful, the target first instruction sequence matching the first instruction sequence exists in the target instruction set; If the match fails or the function types corresponding to the sub-pattern libraries do not correspond to the first function type, the target first instruction sequence matching the first instruction sequence does not exist in the target instruction set.

2. The cross-architecture instruction processing method based on pattern matching and fusion according to claim 1, characterized in that: The sub-pattern library includes one or more of a load-operate-store sub-pattern library, an inter-register operation sub-pattern library, an address calculation and memory access sub-pattern library, and a loop control sub-pattern library; Wherein, the load-operate-store sub-mode library includes the following target first instruction sequences: a load data from memory instruction sequence, an execute ALU instruction sequence, and a store back to memory instruction sequence; The inter-register operation sub-pattern library includes the following target first instruction sequences: a continuous register addition and subtraction operation instruction sequence, a shift instruction sequence and a logic operation instruction sequence; The memory access sub-mode library includes the following target first instruction sequence: an address calculation instruction sequence based on an offset and a base address; The loop control sub-pattern library includes the following target first instruction sequences: a conditional jump operation instruction sequence and a loop instruction sequence.

3. The cross-architecture instruction processing method based on pattern matching and fusion according to claim 1, characterized in that: The method also includes obtaining a first instruction sequence to be processed by: Determine the maximum length of the target first instruction sequence in the target instruction set; A sliding window is used to detect an input first instruction stream, the length of the sliding window is not less than the maximum length, and the first instruction stream includes a plurality of the first instruction sequences.

4. The cross-architecture instruction processing method based on pattern matching and fusion according to claim 3, characterized in that: Two adjacent detection ranges of the sliding window are connected or have overlapping parts.

5. The cross-architecture instruction processing method based on pattern matching and fusion according to claim 1, characterized in that: The first instruction sequence is translated into the second instruction sequence in the following manner: Parallel processing and field separation of multiple instructions in the first instruction sequence, including extracting operation code, operand and prefix information; Each instruction in the first instruction sequence is translated through a lookup table or a preset translation rule, and the second instruction sequence is determined according to the translation results of each instruction.

6. The cross-architecture instruction processing method based on pattern matching and fusion according to claim 5, characterized in that: Translating the first instruction sequence into a second instruction sequence further includes: configuring a cache unit, the cache unit being configured to store a plurality of previous first instruction sequences and a previous second instruction sequence for replacing the previous first instruction sequence, the first instruction set including the previous first instruction sequence, and the second instruction set including the previous second instruction sequence; The usage frequency of the previous first instruction sequence is higher than the preset first usage frequency, and / or the time period of the most recent usage of the previous first instruction sequence relative to the current time period is within a preset time period range; It is determined whether the first instruction sequence is one of the plurality of previous first instruction sequences; if so, the first instruction sequence is directly replaced by the corresponding previous second instruction sequence.

7. The cross-architecture instruction processing method based on pattern matching and fusion according to claim 5, characterized in that: The first instruction sequence is translated into a second instruction sequence, further comprising: The first instruction sequence that does not conform to the specification is identified by checking the legality of the operation code and the operand, and a prompt message is issued.

8. The cross-architecture instruction processing method based on pattern matching and fusion according to claim 1, characterized in that: The method is used to convert a first instruction stream into a second instruction stream, wherein the first instruction stream includes a plurality of first instruction sequences, and the second instruction stream includes a plurality of second instruction sequences.

9. The cross-architecture instruction processing method based on pattern matching and fusion according to claim 8, characterized in that: For a current first instruction sequence in the first instruction stream, if the target first instruction sequence matching the current first instruction sequence does not exist in the target instruction set, translating the current first instruction sequence into a second instruction sequence; When translating the current first instruction sequence, it is simultaneously determined whether there is a target first instruction sequence in the target instruction set that matches the first instruction sequence after the current first instruction sequence.

10. The cross-architecture instruction processing method based on pattern matching and fusion according to claim 1, characterized in that: The first instruction set is based on the X86 architecture, and the second instruction set is based on the RISC-V architecture; or, The first instruction set is based on the ARM architecture, and the second instruction set is based on the RISC-V architecture.

11. The cross-architecture instruction processing method based on pattern matching and fusion according to claim 1, characterized in that: The usage frequency of the target first instruction sequence is higher than a preset second usage frequency.

12. A cross-architecture instruction processor based on pattern matching and fusion, characterized in that: Used to convert an instruction sequence in a first instruction set into an instruction sequence in a second instruction set, wherein the first instruction set and the second instruction set are based on different instruction set architectures; The cross-architecture instruction processor based on pattern matching and fusion includes an instruction fusion hardware module and a translation module; the instruction fusion hardware module is configured to determine whether there is a target first instruction sequence matching the first instruction sequence to be processed in the target instruction set, and if so, replace the first instruction sequence with a target second instruction sequence corresponding to the target first instruction sequence; If not, translating the first instruction sequence into a second instruction sequence using a translation module, where the second instruction sequence includes one or more instructions in a second instruction set; The target instruction set includes a plurality of the target first instruction sequences, the target first instruction sequence includes one or more instructions in the first instruction set, and the target first instruction sequence satisfies the following conditions: The second instruction set includes the target second instruction sequence for replacing the target first instruction sequence, and the number of instructions included in the target second instruction sequence is not greater than the number of instructions included in the corresponding target first instruction sequence; The instruction fusion hardware module is configured with a function mode library, the function mode library includes a plurality of sub-mode libraries, the sub-mode library includes a plurality of the target first instruction sequences with the same function type, and the function types corresponding to the target first instruction sequences included in different sub-mode libraries are different; Each of the sub-pattern libraries is configured with a unique function type, and the function type corresponding to the sub-pattern library corresponds to the function type corresponding to the target first instruction sequence included therein; The instruction fusion hardware module is further configured to determine whether there is a target first instruction sequence matching the first instruction sequence in the target instruction set by: Determine a function type corresponding to the first instruction sequence and use it as a first function type; Determine whether there is a function type corresponding to the sub-pattern library corresponding to the first function type; if so, match the first instruction sequence with a target first instruction sequence included in the sub-pattern library of the corresponding function type; if the match is successful, the target first instruction sequence exists in the target instruction set and matches the first instruction sequence; If the matching fails or the function types corresponding to the sub-pattern libraries do not correspond to the first function type.

13. The cross-architecture instruction processor based on pattern matching and fusion according to claim 12, characterized in that: The translation module includes a hardware translator and a cache unit, wherein the hardware translator is configured to translate the first instruction sequence into a second instruction sequence; The cache unit is configured to store a plurality of previous first instruction sequences and a previous second instruction sequence for replacing the previous first instruction sequence, the first instruction set includes the previous first instruction sequence, and the second instruction set includes the previous second instruction sequence; The usage frequency of the previous first instruction sequence is higher than the preset first usage frequency, and / or the time period of the most recent usage moment of the previous first instruction sequence relative to the current moment is within a preset time period range; The hardware translator is used to translate the first instruction sequence into a second instruction sequence, and it is determined whether the first instruction sequence is one of the multiple prior first instruction sequences. If so, the corresponding prior second instruction sequence is directly used to replace the first instruction sequence; if not, the translation result obtained by translating the first instruction sequence by the hardware translator is used as the second instruction sequence.

Citation Information

Patent Citations

  • Executing programs for a first computer architecture on a computer of a second architecture

    WO2000045257A3