Instruction multiplexing method and device, computer equipment and program product

By modifying the task address of historical memory movement instructions to generate the movement instructions to be executed, the memory consumption problem caused by the dynamic shape sequence of large language models is solved, and the processing efficiency is improved.

CN120929138APending Publication Date: 2025-11-11北京凌川科技有限公司 +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510887316.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing technologies result in excessive memory consumption and negatively impact model execution efficiency when processing dynamic shape sequences of large language models.

Method used

By identifying the associated problem corresponding to the current input problem and modifying the task address in the historical memory transfer instruction to the target address, a transfer instruction to be executed is generated and executed to process the current input problem.

Benefits of technology

It reduces memory usage, improves processing efficiency, and lowers memory space consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120929138A_ABST
    Figure CN120929138A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of deep learning, in particular to an instruction multiplexing method and device, computer equipment and a program product. The method comprises the steps that an associated question corresponding to a current input question is determined, the length of the associated question is consistent with that of the current input question, or the associated question is consistent with a specific sequence of the current input question, and the specific sequence is the sum of the length of the input question and the length of a generated answer; modifying a task address in a historical memory handling instruction into a target address corresponding to the current input problem to obtain a to-be-executed handling instruction, the historical memory handling instruction being a task instruction corresponding to the associated problem; and executing the to-be-executed carrying instruction to process a task corresponding to the current input problem. By adopting the scheme of the invention, the memory space occupied by the large language model can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of deep learning technology, and in particular to an instruction multiplexing method, apparatus, computer device and program product. Background Technology

[0002] Modern deep learning accelerators (such as GPUs or NPUs) are most efficient at operating on tensors with fixed shapes; however, a key characteristic of large language model computation is that it mostly deals with dynamic shapes. Since the compiler cannot know in advance the specific sequence length (input length + generated length) for each inference, it is difficult to generate a single optimal computational graph when faced with different sequence lengths. In other words, the length of the question posed to the large language model varies each time, leading to different shapes and computational flows in the model.

[0003] To handle sequences of varying lengths, several strategies have been employed in related technologies. The first strategy involves padding all sequences to the length of the longest sequence in the current batch, allowing the model to process only fixed shapes. However, this requires the model to perform calculations on invalid padding tokens, and loading and storing these invalid tokens consumes valuable memory bandwidth. The second strategy pre-stores the instructions corresponding to each sequence of length. While this allows for the static processing of all dynamic models, the wide range of sequence lengths results in a large number of instructions needing to be saved, leading to wasted storage space.

[0004] It is evident that the strategies for dealing with dynamic shapes in related technologies require processing a large amount of data, resulting in excessive memory consumption and reduced model execution efficiency.

[0005] Therefore, how to reduce the memory footprint of large language models is an urgent problem to be solved. Summary of the Invention

[0006] Therefore, it is necessary to provide an instruction reuse method, apparatus, computer equipment, and program product that can reduce the memory space occupied by large language models in order to address the above-mentioned technical problems.

[0007] In a first aspect, this application provides an instruction multiplexing method, the method comprising:

[0008] Determine the associated question corresponding to the current input question, wherein the length of the associated question is the same as the length of the current input question, or the associated question is the same as the specific sequence of the current input question, wherein the specific sequence is the sum of the length of the input question and the length of the generated answers;

[0009] The task address in the historical memory transfer instruction is modified to the target address corresponding to the current input problem to obtain the transfer instruction to be executed. The historical memory transfer instruction is the task instruction corresponding to the associated problem.

[0010] Execute the pending transport instructions to process the task corresponding to the current input problem.

[0011] In one embodiment, modifying the task address in the historical memory transfer instruction to the target address corresponding to the current input problem to obtain the transfer instruction to be executed includes:

[0012] Generate and execute control instructions, which are used to write the address offset into the target register corresponding to the historical memory transfer instruction;

[0013] The historical memory transfer instructions are preprocessed to obtain the transfer instructions to be executed; the preprocessing includes reading the address offset from the target register and adding the address offset to the task address corresponding to the memory transfer instruction to obtain the target address.

[0014] In one embodiment, generating and executing control instructions includes:

[0015] Determine the address offset between the task address and the target address;

[0016] A first control instruction is generated based on the address offset, and the first control instruction is executed.

[0017] The first control instruction is used to write the address offset into the target register corresponding to the historical memory transfer instruction.

[0018] In one embodiment, generating and executing control instructions includes:

[0019] Determine the address offset between the task address and the target address, and store the address offset in a temporary address in memory;

[0020] Generate and execute a second control instruction; the second control instruction is used to read the address offset from the temporary address and write the address offset into the target register corresponding to the historical memory transfer instruction.

[0021] In one embodiment, the method further includes:

[0022] When the address offset within the temporary address is exhausted, a third control instruction is generated, which is used to allocate a new address in memory and update the temporary address based on the new address;

[0023] Determine the address offset between the task address and the target address, and write the address offset into the updated temporary address;

[0024] The temporary address contains a preset number of address offsets, and each time it is read, one address offset is consumed.

[0025] In one embodiment, the method further includes:

[0026] The target register is validated. If the target register is found to be valid, the historical memory transfer instructions are preprocessed to obtain the transfer instructions to be executed.

[0027] The preprocessed transport instructions to be executed are marked.

[0028] Secondly, this application also provides an instruction multiplexing apparatus, the apparatus comprising an association problem determination module, a modification module, and an execution module, wherein:

[0029] The associated question determination module is used to determine the associated questions corresponding to the current input question. The length of the associated questions is the same as the length of the current input question, or the associated questions are the same as the specific sequence of the current input question. The specific sequence is the sum of the length of the input question and the length of the generated answers.

[0030] The modification module is used to modify the task address in the historical memory transfer instruction to the target address corresponding to the current input question, so as to obtain the transfer instruction to be executed. The historical memory transfer instruction is the task instruction corresponding to the associated question.

[0031] The execution module is used to execute the transport instructions to process the task corresponding to the current input problem.

[0032] Thirdly, this application also provides a computer device, which includes an instruction preprocessing unit, a memory transfer unit, and a computing unit, wherein:

[0033] The instruction preprocessing unit is used to modify the task address in the historical memory transfer instruction to the target address corresponding to the current input problem, so as to obtain the transfer instruction to be executed. The historical memory transfer instruction is the task instruction corresponding to the associated problem.

[0034] Wherein, the historical memory transfer instruction is the task instruction corresponding to the associated question, the length of the associated question is the same as the length of the current input question, or the specific sequence of the associated question is the same as the specific sequence of the current input question, the specific sequence being the sum of the length of the input question and the length of the generated answer;

[0035] The memory transfer unit is used to receive and execute the transfer instruction to be executed.

[0036] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the instruction multiplexing method as described in any one of the first aspects above.

[0037] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the instruction multiplexing method as described in any one of the first aspects above.

[0038] The aforementioned instruction multiplexing method, apparatus, computer equipment, and program products, for related problems, ensure that the computational process corresponding to the related problem is consistent with the computational process of the current input problem, and therefore the memory movement process in the instruction sequence is also consistent. In other words, by modifying the task address in the historical memory movement instructions corresponding to the related problem to the target address corresponding to the task of the current input problem, the resulting movement instructions to be executed can be applied to the task corresponding to the current input problem. This reduces the number of instructions that need to be loaded into DDR when executing the task corresponding to the current input problem, thereby reducing memory usage and improving the processing efficiency for the current input problem. Attached Figure Description

[0039] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0040] Figure 1 This is a flowchart illustrating an instruction multiplexing method in one embodiment;

[0041] Figure 2 This is a flowchart illustrating the modification of historical memory movement instructions in one embodiment;

[0042] Figure 3 This is a flowchart illustrating the execution of the first control instruction in one embodiment;

[0043] Figure 4 This is a flowchart illustrating the execution of the second control instruction in one embodiment;

[0044] Figure 5 This is a flowchart illustrating the execution of a third control instruction in one embodiment;

[0045] Figure 6This is a structural block diagram of an instruction multiplexing device in one embodiment;

[0046] Figure 7 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0047] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0048] The instruction multiplexing method provided in this application embodiment can be applied to computer devices, wherein the computer devices include memory and processors, and the processor can be an NPU (Neural Network Processing Unit) or a GPU (Graphics Processing Unit), that is, a processor device suitable for performing computational tasks of large language models.

[0049] In one exemplary embodiment, such as Figure 1 As shown, an instruction multiplexing method is provided. Taking the application of this method to an NPU as an example, the method includes the following steps 110-130, wherein:

[0050] Step 110: Determine the associated question corresponding to the current input question. The length of the associated question is the same as the length of the current input question, or the associated question is the same as the specific sequence of the current input question. The specific sequence is the sum of the length of the input question and the length of the generated answer.

[0051] In this embodiment of the application, in related technologies, when the NPU executes a task corresponding to each input question, it needs to load the instruction sequence corresponding to that task into DDR (Double Data Rate) memory. The historical instruction sequences corresponding to tasks that have been processed within a historical period are not directly deleted, but are retained in DDR for a preset duration. The duration of both the historical period and the preset duration can be preset; this embodiment does not impose specific limitations and can be set by the user. The historical instruction sequence includes historical memory movement instructions and historical calculation instructions.

[0052] For the current input problem, before starting execution of the corresponding task, it is pre-determined whether a historical input problem of the same length exists within the historical period. If such a historical input problem exists, it is identified as a related problem of the current input problem. When processing the task corresponding to the current input problem, the historical memory transfer instructions in the historical instruction sequence corresponding to the related problem in DDR can be directly retrieved. If it is determined that no related problem exists before starting execution of the corresponding task, the instruction sequence corresponding to the current input problem is directly invoked to begin execution.

[0053] Furthermore, during the execution of the task corresponding to the current input question and the generation of the answer, it is determined whether there is a historical input question with a specific sequence that is consistent with the current input question within the historical period; if so, the historical question is identified as a related question. For the historical instruction sequence corresponding to the related question, the historical memory transfer instruction corresponding to the point where the specific sequence length of the related question is consistent with the current input question is retrieved.

[0054] Step 120: Modify the task address in the historical memory transfer instruction to the target address corresponding to the current input problem to obtain the transfer instruction to be executed. The historical memory transfer instruction is the task instruction corresponding to the associated problem.

[0055] In the embodiments of this application, memory transfer instructions are used to transfer data from DDR to the computing unit, or to transfer data from the computing unit to DDR. Therefore, the task address corresponding to the historical memory transfer instructions is the address corresponding to the historical input problem. If the historical memory transfer instructions are to be used to execute the task corresponding to the current input problem, the memory transfer instructions need to be able to point to the target address corresponding to the task of the current input problem. That is, the memory transfer instructions need to be processed to modify their corresponding task address.

[0056] Step 130: Execute the pending transfer instructions to process the task corresponding to the current input problem.

[0057] In this embodiment of the application, after modifying the task address in the historical memory transfer instruction to the target address corresponding to the task of the current input problem, the result is a transfer instruction to be executed; the transfer instruction to be executed can be used to execute the task corresponding to the current input problem; when executing the transfer instruction to be executed, data is transferred from DDR to the computing unit, or data is transferred from the computing unit to DDR.

[0058] In the above instruction reuse method, for the associated problem, the calculation process corresponding to the associated problem is consistent with the calculation process of the current input problem, and therefore the memory movement process in the instruction sequence is also consistent. That is to say, after modifying the task address in the historical memory movement instruction corresponding to the associated problem to the target address corresponding to the task of the current input problem, the resulting movement instruction to be executed can be applied to the task corresponding to the current input problem, thereby reducing the number of instructions that need to be loaded into DDR when executing the task corresponding to the current input problem, thereby reducing the memory occupation and thus improving the processing efficiency for the current input problem.

[0059] Furthermore, this section explains the historical instruction sequence of related questions, which can be applied to historical memory-moving instructions related to the current input question. For related questions whose length matches the current input question, all historical memory-moving instructions following the start of the historical instruction sequence of the related question can be applied to the current input question (until the task corresponding to the current input question ends, or all historical memory-moving instructions are reused). For example, if the length of a historical question is 50 and the answer length is 100, its corresponding historical instruction sequence contains 150 instructions. If the length of the current input question is 50, then when executing the task for the current input question, reuse can begin from the first historical memory-moving instruction in the historical instruction sequence of the related questions. During task execution, if it is determined that the answer length of the current input question is 50, then the task corresponding to the current input question can end before all historical memory-moving instructions in the historical instruction sequence are fully reused. In this process, some historical memory-moving instructions corresponding to related questions are reused.

[0060] For example, if a historical question has a length of 50 and an answer length of 100, its corresponding historical instruction sequence contains 150 instructions. If the current input question has a length of 50, then when executing the task for the current input question, reuse can begin from the first historical memory movement instruction in the historical instruction sequence of the associated questions. During task execution, if it is determined that the answer length of the current input question is 200, then after fully reusing the historical memory movement instructions in the historical instruction sequence, the instructions corresponding to the task for the current input question are further loaded into DDR to continue execution, thus ending the task for the current input question. In this process, the historical memory movement instructions corresponding to the associated questions are fully reused.

[0061] For related questions where the length of the specific sequence matches that of the current input question, historical memory-moving instructions following the point of consistency in the historical instruction sequence of the related question can be applied to the current input question (until the task corresponding to the current input question ends, or when all historical memory-moving instructions are reused). For example, if the length of the historical question is 50 and the answer length is 100, there are 130 instructions in its corresponding historical instruction sequence. If the length of the current input question is 30, then when executing the task for the current input question, when the input length is 20, reuse can begin from the first historical memory-moving instruction in the historical instruction sequence of the related question; during task execution, if it is determined that the final answer length of the task corresponding to the current input question is 200, then when the answer length corresponding to the current input question is 120, all historical memory-moving instructions in the historical instruction sequence can be fully reused.

[0062] In one exemplary embodiment, such as Figure 2 As shown, step 120 may specifically include steps 121 and 122, wherein:

[0063] Step 121: Generate and execute control instructions. These control instructions are used to write the address offset into the target register corresponding to the historical memory transfer instruction.

[0064] Step 122: Preprocess the historical memory transfer instructions to obtain the transfer instructions to be executed; the preprocessing includes reading the address offset from the target register and adding the address offset to the corresponding task address in the memory transfer instruction to obtain the target address.

[0065] Specifically, the NPU includes an instruction preprocessing unit, a memory transport unit, and a control instruction execution module. After the NPU generates control instructions, the control instruction execution module executes these instructions, writing the address offset into the target register corresponding to the historical memory transport instruction. Further, the preprocessing module preprocesses the historical memory transport instructions by reading the address offset from the target register corresponding to the historical memory transport instruction, adding the address offset to the task address corresponding to the historical memory transport instruction to obtain the target address, and modifying the task address in the historical memory transport instruction to the target address, resulting in the transport instruction to be executed. Finally, the instruction preprocessing unit sends the transport instruction to be executed to the memory transport unit for execution, enabling data to be transported between the target address and the computation unit. The address offset is the offset of the target address relative to the task address in the historical memory transport instruction.

[0066] Furthermore, the memory movement instructions are analyzed in detail. These instructions include memory movement instruction 1 and memory movement instruction 2.

[0067] For memory transfer instruction 1, the instruction name is load, which means transferring instructions from DDR to the NPU CORE (NPU core, also known as the computing unit). The instruction format of memory transfer instruction 1 is "load ddr_addr, size, reg". Here, addr is the address flag, ddr indicates the location (address) of the data to be transferred in DDR, size indicates the quantity to be transferred, and reg is the register flag; reg takes a value from 0 to 7. When reg3 is used in the instruction, it indicates that the register with the label 3 is the target register of this instruction.

[0068] For memory transfer instruction 2, the instruction name is store, which means moving data from the NPU CORE (NPU core, also known as the computing unit) to DDR. The instruction format of memory transfer instruction 2 is store ddr_addr, size, reg, where addr is the address flag, ddr indicates the location (address) of the data to be moved in DDR, size indicates the quantity to be moved, and reg is the register flag; reg takes a value from 0 to 7. When reg3 is used in the instruction, it indicates that the register with the label 3 is the target register of this instruction.

[0069] The instruction preprocessing unit preprocesses a historical memory transfer instruction (memory transfer instruction 1 or memory transfer instruction 2) by adding the value (address offset) in the corresponding label register (target register) inside the memory transfer instruction to ddr_addr, thereby modifying ddr in the historical memory transfer instruction to the target address, obtaining the transfer instruction to be executed, and sending the transfer instruction to be executed to the memory transfer unit.

[0070] Furthermore, when processing instructions, the instruction preprocessing unit cannot distinguish whether a memory transfer instruction needs modification or is an already modified instruction to be executed. To improve the accuracy of instruction processing, each modified memory transfer instruction is marked as an instruction to be executed. This marking can be done by changing the register labels in the instruction to be executed to invalid labels. For example, if the register labels range from 0 to 7, the register labels in the instruction to be executed can be changed to 8. Since there is no register labeled 8, the register labels in the instruction to be executed are invalid.

[0071] Furthermore, after the instruction preprocessing unit receives a memory transfer instruction, it first verifies the validity of the destination register, that is, determines the validity of the register label in the instruction. If the destination register is determined to be valid, the current instruction is identified as a historical memory transfer instruction that needs preprocessing, and the historical memory transfer instruction is preprocessed to obtain the transfer instruction to be executed. If the destination register in the instruction is determined to be invalid, the instruction is determined to be a memory transfer instruction to be executed. In this case, the instruction preprocessing unit does not perform any processing on the instruction, but directly sends the instruction to the memory transfer unit for execution.

[0072] Furthermore, step 121 can be implemented in two ways, which will be described below.

[0073] The first implementation of step 121 is as follows: determine the address offset between the task address and the target address; generate a first control instruction based on the address offset and execute the first control instruction; the first control instruction is used to write the address offset into the target register corresponding to the historical memory transfer instruction.

[0074] In related technologies, after the current input problem is compiled into characters by the upper-layer software / compiler, a computational flow for the current input problem is generated based on the characters. Furthermore, the NPU loads the corresponding instructions into DDR for the computational flow and executes them. In other words, when the computational flow corresponding to the current input problem is determined, the target address corresponding to the current input problem has already been determined.

[0075] Specifically, for historical memory transfer instructions, the corresponding task address is specified in the instruction, and the target address corresponding to the current input problem has been determined; therefore, the offset of the target address relative to the task address in the historical memory transfer instructions can be directly determined.

[0076] Furthermore, the instruction name of the first control instruction is rset1, and the instruction format is rset1 value, reg. The first control instruction is also executed by the instruction preprocessing unit. When the instruction preprocessing unit executes the first control instruction, it fills the value in the first control instruction into the register (target register) corresponding to the label in the instruction. The address offset can be directly written to the target register through the first control instruction, allowing the instruction control unit to directly retrieve the address offset from the target register, thereby modifying historical memory transfer instructions. Writing the address offset to the target register through the first control instruction is a highly efficient method.

[0077] The second implementation of step 121 is as follows: determine the address offset between the task address and the target address, and store the address offset in a temporary address in memory; generate and execute a second control instruction; the second control instruction is used to read the address offset from the temporary address and write the address offset into the target register corresponding to the historical memory transfer instruction.

[0078] Specifically, a preset number of address offsets are written into the temporary address, and one address offset is consumed each time it is read. The NPU can execute a certain number of memory transfer instructions per unit cycle; however, when the current input problem requires too many reusable instructions, it is not possible to modify all historical memory transfer instructions at once. Therefore, in the second implementation, the address offsets need to be pre-stored in DDR, and subsequently, the historical memory transfer instructions that need to be reused are modified in real time by reading the address offsets stored in DDR.

[0079] The memory scheduling system allocates a temporary address in DDR to store a defined address offset, and writes the address offset into this temporary address, with the number of offsets written being a preset number. This preset number can be predetermined or depend on the available storage space of the temporary address; this embodiment does not impose a specific limitation on it.

[0080] Furthermore, the second control instruction is named rset2, and its format is "rset2 ddr_addr, reg". This second control instruction is also executed by the instruction preprocessing unit. It indicates that the instruction preprocessing unit adds the address (temporary address) and the start address (start_addr) in the instruction, reads a number (address offset) from the target address of the DDR according to the sum, and fills it into its corresponding label register (target register).

[0081] The starting address is the reference address of all instructions at the start of the task, which can also be used as a zero address; the starting address is allocated by the memory import system.

[0082] Furthermore, when the address offset within the temporary address is exhausted, a third control instruction is generated. This third control instruction is used to allocate a new address in memory and update the temporary address based on the new address.

[0083] Specifically, when the address offset stored in the initially allocated temporary address is exhausted, a new address is reallocated in DDR to store the address offset, and the temporary address is updated based on this new address. Furthermore, the starting address of the global reference is updated. The address offset between the task address and the target address is determined, and this offset is written to the updated temporary address.

[0084] The third control instruction is named rset3 and has the format "rset3 addr". This instruction is executed by the instruction preprocessing unit, which updates the starting address (start_addr) to the address specified in this instruction (the new address). This ensures that all subsequent reused historical memory movement instructions can accurately obtain the address offset.

[0085] The logic of the above control commands being executed is explained in detail below:

[0086] Reference Figure 3 This is a schematic diagram of the execution flow of the first control instruction (control instruction 1). After the control instruction 1 is passed to the instruction preprocessing unit, the instruction preprocessing unit rewrites the value of the register corresponding to reg in the instruction to value (address offset), and then ends.

[0087] Reference Figure 4 Here is a schematic diagram of the execution flow of the second control instruction (control instruction 2): The instruction preprocessing unit adds the address (ddr_addr) and the start address (start_addr) in the instruction, reads a number from DDR according to the address obtained by the addition, and fills it into its corresponding label register (reg).

[0088] Reference Figure 5 Here is a schematic diagram of the execution flow of the third control instruction (control instruction 3): The instruction preprocessing unit updates the starting address (start_addr) to the address (addr) in the instruction.

[0089] By executing the first, second, third, and second control instructions, the address offset can be written to the target register. This allows the task corresponding to the current input question to read the address offset from the target register and modify the historical memory transfer instructions that need to be reused. The instruction reuse method provided in this application can statically process the dynamic shape of large language models caused by varying request lengths (input question lengths), and reduce the storage space occupied by instructions, thereby improving network execution efficiency. Furthermore, the static instruction storage space is reduced to less than 10% of the original size.

[0090] In fact, the technical means of this application make it possible to statically convert a large language model completely when the request length is unknown. Depending on the network structure, the instruction storage space after staticization is reduced to 1 / 32 to 1 / 64 of the original. Furthermore, dynamically calculating and optimizing the network shape based on the request length is time-consuming; therefore, under real-time requirements, dynamic calculation and optimization methods may not be fully feasible. However, through the instruction reuse method of this application, the network can be pre-staticized. Since there are no real-time requirements, all optimization methods can be used, resulting in better network performance.

[0091] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0092] Based on the same inventive concept, this application also provides an instruction multiplexing apparatus for implementing the instruction multiplexing method described above. The solution provided by this apparatus is similar to the implementation described in the above method; therefore, the specific limitations in one or more instruction multiplexing apparatus embodiments provided below can be found in the limitations of the instruction multiplexing method described above, and will not be repeated here.

[0093] In one exemplary embodiment, such as Figure 6 As shown, an instruction multiplexing device 600 is provided, including an association problem determination module 601, a modification module 602, and an execution module 603, wherein:

[0094] The associated question determination module 601 is used to determine the associated questions corresponding to the current input question. The length of the associated questions is the same as the length of the current input question, or the associated questions are the same as the specific sequence of the current input question. The specific sequence is the sum of the length of the input question and the length of the generated answers.

[0095] Modification module 602 is used to modify the task address in the historical memory transfer instruction to the target address corresponding to the current input problem, so as to obtain the transfer instruction to be executed. The historical memory transfer instruction is the task instruction corresponding to the associated problem.

[0096] The execution module 603 is used to execute the transfer instructions to process the task corresponding to the current input problem.

[0097] In one embodiment, the modification module 602 is specifically used for:

[0098] Generate and execute control instructions, which are used to write the address offset into the target register corresponding to the historical memory transfer instruction;

[0099] The historical memory transfer instructions are preprocessed to obtain the transfer instructions to be executed. The preprocessing includes reading the address offset from the target register and adding the address offset to the corresponding task address in the memory transfer instruction to obtain the target address.

[0100] In one embodiment, the modification module 602 is specifically used for:

[0101] Determine the address offset between the task address and the target address;

[0102] The first control instruction is generated based on the address offset, and then executed.

[0103] The first control instruction is used to write the address offset into the target register corresponding to the historical memory transfer instruction.

[0104] In one embodiment, the modification module 602 is specifically used for:

[0105] Determine the address offset between the task address and the target address, and store the address offset in a temporary address in memory;

[0106] Generate and execute a second control instruction; the second control instruction is used to read the address offset from the temporary address and write the address offset into the target register corresponding to the historical memory transfer instruction.

[0107] In one embodiment, the instruction multiplexing device 600 further includes an update module, specifically used for:

[0108] When the address offset within the temporary address is exhausted, a third control instruction is generated. This third control instruction is used to allocate a new address in memory and update the temporary address based on the new address.

[0109] Determine the address offset between the task address and the target address, and write the address offset to the updated temporary address;

[0110] The temporary address contains a preset number of address offsets, and each time it is read, one address offset is consumed.

[0111] In one embodiment, the instruction multiplexing device 600 further includes a verification module, specifically used for:

[0112] The target register is validated. If the target register is found to be valid, the historical memory transfer instructions are preprocessed to obtain the transfer instructions to be executed.

[0113] The preprocessed transport instructions to be executed are marked.

[0114] Each module in the aforementioned instruction multiplexing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0115] In one exemplary embodiment, a computer device is provided, which may be a CPU or an NPU, and its internal structure diagram may be as follows. Figure 7 As shown. The computer device includes an instruction preprocessing unit and a memory handling unit, wherein:

[0116] The instruction preprocessing unit (containing a set of labeled registers and a starting address) is used to modify the task address in the historical memory transfer instruction to the target address corresponding to the current input problem, so as to obtain the transfer instruction to be executed. The historical memory transfer instruction is the task instruction corresponding to the associated problem.

[0117] Among them, the historical memory transfer instruction is the task instruction corresponding to the associated question. The length of the associated question is the same as the length of the current input question, or the specific sequence of the associated question is the same as the current input question. The specific sequence is the sum of the length of the input question and the length of the generated answer.

[0118] The memory transfer unit is used to receive and execute transfer instructions to be executed.

[0119] Furthermore, the computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and databases. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements an instruction multiplexing method.

[0120] Those skilled in the art will understand that Figure 7 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0121] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the instruction multiplexing methods described in the above-described instructions multiplexing method embodiments.

[0122] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of any of the instruction multiplexing methods described in the above-described instructions multiplexing method embodiments.

[0123] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0124] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0125] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0126] The above embodiments are merely illustrative of several implementation methods of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. An instruction multiplexing method, characterized in that, The method includes: Determine the associated question corresponding to the current input question. The length of the associated question is the same as the length of the current input question, or the associated question is the same as the specific sequence of the current input question. The specific sequence is the sum of the length of the input question and the length of the generated answers. The task address in the historical memory transfer instruction is modified to the target address corresponding to the current input question to obtain the transfer instruction to be executed. The historical memory transfer instruction is the task instruction corresponding to the associated question. Execute the pending transport instructions to process the task corresponding to the current input problem.

2. The method according to claim 1, characterized in that, The step of modifying the task address in the historical memory transfer instruction to the target address corresponding to the current input problem, to obtain the transfer instruction to be executed, includes: Generate and execute control instructions, which are used to write the address offset into the target register corresponding to the historical memory transfer instruction; The historical memory transfer instructions are preprocessed to obtain the transfer instructions to be executed; the preprocessing includes reading the address offset from the target register and adding the address offset to the task address corresponding to the memory transfer instruction to obtain the target address.

3. The method according to claim 2, characterized in that, The generation and execution of control instructions includes: Determine the address offset between the task address and the target address; A first control instruction is generated based on the address offset, and the first control instruction is executed. The first control instruction is used to write the address offset into the target register corresponding to the historical memory transfer instruction.

4. The method according to claim 2, characterized in that, The generation and execution of control instructions includes: Determine the address offset between the task address and the target address, and store the address offset in a temporary address in memory; Generate and execute a second control instruction; the second control instruction is used to read the address offset from the temporary address and write the address offset into the target register corresponding to the historical memory transfer instruction.

5. The method according to claim 4, characterized in that, The method further includes: When the address offset within the temporary address is exhausted, a third control instruction is generated, which is used to allocate a new address in memory and update the temporary address based on the new address; Determine the address offset between the task address and the target address, and write the address offset into the updated temporary address; The temporary address contains a preset number of address offsets, and each time it is read, one address offset is consumed.

6. The method according to any one of claims 2 to 5, characterized in that, The method further includes: The target register is validated. If the target register is found to be valid, the historical memory transfer instructions are preprocessed to obtain the transfer instructions to be executed. The preprocessed transport instructions to be executed are marked.

7. An instruction multiplexing device, characterized in that, The device includes an association problem determination module, a modification module, and an execution module, wherein: The associated question determination module is used to determine the associated questions corresponding to the current input question. The length of the associated questions is the same as the length of the current input question, or the associated questions are the same as the specific sequence of the current input question. The specific sequence is the sum of the length of the input question and the length of the generated answers. The modification module is used to modify the task address in the historical memory transfer instruction to the target address corresponding to the current input question, so as to obtain the transfer instruction to be executed. The historical memory transfer instruction is the task instruction corresponding to the associated question. The execution module is used to execute the transport instructions to process the task corresponding to the current input problem.

8. A computer device, characterized in that, The computer device includes an instruction preprocessing unit and a memory transfer unit, wherein: The instruction preprocessing unit is used to modify the task address in the historical memory transfer instruction to the target address corresponding to the current input problem, so as to obtain the transfer instruction to be executed. The historical memory transfer instruction is the task instruction corresponding to the associated problem. Wherein, the historical memory transfer instruction is the task instruction corresponding to the associated question, the length of the associated question is the same as the length of the current input question, or the specific sequence of the associated question is the same as the specific sequence of the current input question, the specific sequence being the sum of the length of the input question and the length of the generated answer; The memory transfer unit is used to receive and execute the transfer instruction to be executed.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.