Thread migration method and device

By generating and sending instructions containing context information in a multi-core system, the problem of CPU being unable to perform speculative execution after thread migration is solved, and thread execution efficiency and processor performance are improved.

CN120631518APending Publication Date: 2025-09-12HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410278747.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-12
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

In a multi-core system, during thread migration, speculative execution cannot be implemented because the CPU after migration does not store the context information corresponding to the thread, which affects the thread execution rate.

Method used

A second instruction containing context information is generated by the first processor and sent to the second processor so that the second processor can execute the thread, including adding context information in a modifiable field or using a pre-stored correspondence set to select an appropriate instruction to ensure consistency of instruction content and speculative execution.

Benefits of technology

The execution efficiency of threads in the migrated CPU is improved, speculative execution is achieved, and processor performance is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120631518A_ABST
    Figure CN120631518A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a thread migration method and device.The method comprises the steps that first context information generated in the process of executing a first thread is stored in a first processor, and in the process that the first thread is migrated from the first processor to a second processor, the first processor executes the first context information according to a first instruction and the first context information; generating a second instruction; the first instruction is an instruction required for executing the first thread; the first processor sends a second instruction to the second processor; the second instruction is used for supporting the second processor to execute the first thread; the second processor determines first context information corresponding to the first thread according to the second instruction; the first context information is used for the second processor to execute the first thread. Thus, the second processor obtains the first context information, that is, the first thread can be executed based on the first context information, and the execution efficiency of the thread in the migrated CPU can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of communications, and in particular to a thread migration method and device. Background Art

[0002] In a multi-core system (i.e., multiple central processing units (CPUs)), each CPU core runs multiple threads, and each thread generates multiple contexts during execution. Speculative execution (or speculative execution) is a feature of the CPU. Common speculative executions include prefetching and branch prediction. Speculative execution occurs when the CPU predicts the execution dependencies of an instruction in advance based on context information (i.e., context information generated by the previous CPU) when executing the instruction a second time. This reduces pipeline stalls caused by dependencies and improves processor performance. In other words, the context information corresponding to the thread can be used to predict the execution dependencies of the instruction.

[0003] With system iterations and updates, the number of CPU cores in multi-core systems has gradually increased, for example, to hundreds or even thousands of cores. Accordingly, the number of threads in these multi-core systems has doubled. In scenarios where the number of business threads far exceeds the number of CPU cores (i.e., oversubscription), threads frequently migrate between different CPU cores. These migrations typically only include the thread's corresponding instructions and the data required to execute them. The migrated CPUs do not store the thread's contextual information, making speculative execution impossible. Therefore, the thread migration process affects the thread's execution rate in the migrated CPUs. Summary of the Invention

[0004] Embodiments of the present application provide a thread migration method and apparatus for improving the execution efficiency of threads in a CPU after migration.

[0005] In a first aspect, the present application provides a thread migration method, which is applied to a first processor, or a component (such as a processor, chip, chip system, circuit, or other) in the first processor, or a software module. Taking the method applied to the first processor as an example, the method may include: during the migration of a first thread from the first processor to a second processor, the first processor generates a second instruction based on a first instruction and first context information; the first instruction is an instruction required to execute the first thread; the first processor sends the second instruction to the second processor; and the second instruction is used to support the second processor in executing the first thread.

[0006] Using this method, the first processor can send the second instruction to the second processor, and the second processor can obtain the first context information according to the second instruction, that is, it can execute the first thread based on the first context information, which can improve the execution efficiency of the thread in the migrated CPU.

[0007] In one possible design, the first instruction includes a modifiable field, and the first processor can add the first context information to the modifiable field of the first instruction to obtain the second instruction.

[0008] In a possible design, the first instruction and the second instruction both include an instruction identifier and instruction content; the instruction identifiers of the first instruction and the second instruction are different, and the instruction contents of the first instruction and the second instruction are the same.

[0009] In one possible design, the first processor may obtain a first correspondence set between multiple pre-stored instructions and multiple context information; the first processor may select a second instruction from multiple instructions in the first correspondence set based on the first instruction and the first context information.

[0010] In one possible design, the first processor can determine at least one alternative instruction that has a corresponding relationship with the first context information from multiple instructions in the first correspondence set; the first processor can select an alternative instruction with the same instruction content as the first instruction from at least one alternative instruction as the second instruction.

[0011] In one possible design, the first instruction includes a modifiable mapping identifier field, and the first processor can obtain a second set of correspondences between multiple pre-stored mapping identifiers and multiple context information; the first processor can determine the target mapping identifier corresponding to the first context information in the second correspondence set; the first processor can add the target mapping identifier in the mapping identifier field of the first instruction to obtain a second instruction.

[0012] In one possible design, the first processor may send the second instruction to the second processor through a coherent cache device.

[0013] In one possible design, the first processor may obtain the first instruction through a coherent cache device.

[0014] In one possible design, the first processor may obtain second context information; the second context information is different from the first context information; the first processor may update the second instruction based on the second context information, and the first processor may also send the updated second instruction to the second processor.

[0015] In one possible design, the first processor and the second processor are set in the same device.

[0016] In a second aspect, the present application provides a thread migration method, which is applied to a second processor, or a component (such as a processor, chip, chip system, circuit, or other, etc.) in the second processor, or a software module. Taking the method applied to the second processor as an example, the method may include: the second processor receiving a second instruction from the first processor during the process of migrating the first thread from the first processor to the second processor; the second processor determining first context information corresponding to the first thread based on the second instruction; and the first context information being used by the second processor to execute the first thread.

[0017] In one possible design, the communication unit 601 is specifically configured to receive the second instruction from the first processor through a coherent cache device.

[0018] In a third aspect, the present application further provides a thread migration device. The thread migration device can execute the method described in the first or second aspects above, or the solutions in each possible design. The thread migration device can be a chip or circuit capable of executing the functions corresponding to the above methods, or a device including such a chip or circuit.

[0019] In one possible design, the thread migration device includes a communication unit for receiving and / or sending data; the thread migration device also includes a processing unit for implementing the method described in any one of the possible designs of the first or second aspects. The aforementioned functions can be implemented via hardware or by hardware executing corresponding software. The hardware or software includes one or more units corresponding to the aforementioned functions.

[0020] In a fourth aspect, the present application further provides a thread migration device. The thread migration device can execute the method described in the first or second aspects above, or the solutions in each possible design. The thread migration device includes a processor. When the processor executes instructions, the thread migration device, or a device equipped with the thread migration device, executes the method described in any of the possible designs described in the first or second aspects above.

[0021] Optionally, the thread migration device may further include a memory for storing computer-executable program code, and the program agent may include the aforementioned instructions. The memory may be located inside or outside the thread migration device, which is not limited in this application. The memory may be coupled to the processor.

[0022] The thread migration device may further include a communication interface. Optionally, if the thread migration device is a chip or a circuit, the communication interface may be an input / output interface of the chip, such as an input / output pin.

[0023] In a fifth aspect, the present application provides a system comprising at least one of the following: a first processor that executes the method of the first aspect or a second processor that executes the method of the second aspect.

[0024] Optionally, the system may further include a consistent cache device.

[0025] In a sixth aspect, the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a computer, the computer executes a method in any possible design shown in the first or second aspect above.

[0026] In the seventh aspect, the present application provides a computer program product, in which a computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are called by a computer, the computer executes the method in any possible design shown in the first aspect or the second aspect above.

[0027] In an eighth aspect, the present application provides a chip, comprising a processor configured to execute the method of any one of the possible designs described in the first or second aspects above. Optionally, the chip may further comprise a communication interface configured to input and / or output signaling or data. Optionally, the chip may further comprise a memory configured to store the aforementioned computer program; the processor is coupled to the memory, and the processor may read the computer program stored in the memory to execute the method of any one of the possible designs described in the first or second aspects above.

[0028] In addition, the technical effects brought about by the second to eighth aspects can be found in the description of each possible solution in the first aspect above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 A schematic diagram of the architecture of a multi-core system provided in an embodiment of the present application;

[0030] Figure 2 A flowchart of a thread migration method provided in an embodiment of the present application;

[0031] Figure 3 An example diagram of a thread migration method provided in an embodiment of the present application;

[0032] Figure 4 An example diagram of another thread migration method provided in an embodiment of the present application;

[0033] Figure 5 An example diagram of another thread migration method provided in an embodiment of the present application;

[0034] Figure 6A schematic structural diagram of a thread migration device provided in an embodiment of the present application;

[0035] Figure 7 A schematic diagram of the structure of another thread migration device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0036] In order to make the purpose, technical solutions and beneficial effects of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0037] In the description of this application, unless otherwise specified, “ / ” means or, for example, A / B can mean A or B; “and / or” in this application is merely a description of the association relationship of associated objects, indicating that three relationships can exist, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of this application, “at least one” means one or more items, and “multiple items” means two or more items. In the description of this application, words such as “first” and “second” are only used for the purpose of distinguishing the description, and cannot be understood as indicating or implying relative importance, nor can they be understood as indicating or implying order.

[0038] Figure 1 Schematic diagram of a multi-core system applicable to an embodiment of the present application, wherein the system includes multiple CPUs (e.g., N+1), which are respectively identified as CPU 0, CPU 1, ..., CPU N. The multiple CPUs can share the same cache space (i.e., coherent cache).

[0039] In some instances, thread migration can be performed between any two CPUs (via a coherent cache). For example, CPU 0 can store instructions corresponding to a thread in a coherent cache, and CPU N can read instructions corresponding to the thread from the coherent cache.

[0040] In order to improve the execution efficiency of threads in the CPU after migration, the embodiment of the present application provides a thread migration method. Figure 1 The multi-core system shown is implemented.

[0041] The following will introduce the thread migration method provided in the embodiment of the present application with reference to the accompanying drawings. Figure 2 A thread migration method provided in an embodiment of the present application may include the following steps:

[0042] S201: During migration of the first thread from the first processor to the second processor, the first processor generates a second instruction based on the first instruction and first context information corresponding to the first thread; wherein the first processor stores the first context information generated during execution of the first thread; and the first instruction is an instruction required for executing the first thread.

[0043] In the embodiment of the present application, a thread may refer to a software thread.

[0044] Optionally, before executing S201, the first processor may obtain the first instruction while executing the first thread. For example, the first processor may obtain the first instruction through a coherent cache. The first context information is context information generated when the first thread executes the instruction corresponding to the first thread.

[0045] In one possible design, the first instruction includes a modifiable field, and the method for the first processor to generate the second instruction may include: the first processor adds first context information to the modifiable field of the first instruction to obtain the second instruction.

[0046] In another possible design, the first instruction and the second instruction both include an instruction identifier and instruction content; the instruction identifiers of the first instruction and the second instruction are different, and the instruction contents of the first instruction and the second instruction are the same.

[0047] In some embodiments, the method for the first processor to generate the second instruction may include: the first processor obtains a first correspondence set between multiple pre-stored instructions and multiple context information; the first processor selects the second instruction from multiple instructions in the first correspondence set based on the first instruction and the first context information.

[0048] Optionally, the method for the first processor to select the second instruction from multiple instructions includes: the first processor determines at least one alternative instruction that has a correspondence with the first context information from multiple instructions in the first correspondence set; the first processor selects an alternative instruction with the same instruction content as the first instruction from at least one alternative instruction as the second instruction.

[0049] In another possible design, the first instruction includes a modifiable mapping identifier field, and the method for the first processor to generate the second instruction may include: the first processor obtains a second set of correspondences between multiple pre-stored mapping identifiers and multiple context information; the first processor determines the target mapping identifier corresponding to the first context information in the second correspondence set; the first processor adds the target mapping identifier in the mapping identifier field of the first instruction to obtain the second instruction.

[0050] S202: The first processor sends a second instruction to the second processor; in response, the second processor receives the second instruction from the first processor, wherein the second instruction is used to support the second processor in executing the first thread.

[0051] Optionally, the first processor and the second processor may be in the same device. Figure 1 The multi-core system shown includes a first processor (eg, CPU 0) and a second processor (eg, CPU N).

[0052] In some examples, the process of the first processor sending the second instruction to the second processor includes: the first processor obtains second context information; the second context information is different from the first context information; the first processor updates the second instruction based on the second context information, and sends the updated second instruction to the second processor.

[0053] In one possible design, a first processor sends a second instruction to a second processor via a coherent cache device; correspondingly, the second processor receives the second instruction from the first processor via the coherent cache device. For example, the first processor stores the second instruction in the coherent cache, and the second processor retrieves the second instruction from the coherent cache.

[0054] S203: The second processor determines first context information corresponding to the first thread according to the second instruction; the first context information is used by the second processor to execute the first thread.

[0055] In some examples, the second processor can perform speculative execution based on the first context information to execute the first thread. Speculative execution can improve the execution efficiency of the first thread.

[0056] like Figure 3 As shown, CPU A includes at least one micro-architecture module A and a modification module A; CPU B includes at least one micro-architecture module B and a modification module B. Based on the above Figure 2 The thread migration method shown below is combined with Figure 3 , an example of thread migration between CPU A and CPU B is described. The thread migration process may include the following steps:

[0057] S301: Pre-embed instruction 1 in the coherent cache; the instruction 1 is the instruction required for executing thread A.

[0058] In some examples, instruction 1 includes R+1 fields, identified as field 0 to field R respectively; instruction 1 can exist in the following three forms (an instruction that satisfies any of the following forms is also referred to as a modifiable instruction, BDcond):

[0059] Form 1:

[0060] As shown in Table 1, the i-th field to the i+m-th field of instruction 1 are called reserved fields. The reserved fields can be read and written by the CPU, and the other fields are used to implement the content defined by the instruction set architecture (ISA). For example, instruction 1 may include the following fields:

[0061]

[0062] Table 1

[0063] When the form of instruction 1 is form 1, any CPU can read and write the contents of the reserved field in instruction 1.

[0064] Optionally, the field where the reserved domain segment is located can be in any m+1 fields of instruction 1. The above Table 1 is only an example and does not constitute a limitation to this application.

[0065] Optionally, the reserved domain segment can be m continuous fields or m discontinuous fields, which is not limited in this application.

[0066] Optionally, when the form of instruction 1 is form 1, instruction 1 may further include a modifiable instruction flag; wherein the modifiable instruction flag is used to indicate whether the instruction has a reserved field. As shown in Table 2, one implementation of instruction 1 includes the following fields:

[0067]

[0068]

[0069] Table 2

[0070] Among them, fields 0 to 3 are used to implement some ISA-defined contents (such as conditional coding); field 4 is the aforementioned reserved field segment; field 5 to field 23 are used to implement some ISA-defined contents (such as immediate values), and immediate values ​​are used to indicate the jump range corresponding to instruction 1; field 24 is the aforementioned modifiable instruction flag; field 25 to field 31 are used to implement some ISA-defined contents (such as constants), and constants are used to indicate the direct conditional jump flag corresponding to instruction 1.

[0071] Form 2:

[0072] As shown in Table 3, the 0th to sth fields store instruction identifiers, the Rjth to Rth fields store modifiable instruction identifiers (flags), and the other fields are used to implement the contents defined by the ISA. For example, instruction 1 may include the following fields:

[0073]

[0074] Table 3

[0075] The modifiable instruction flag is used to indicate that the instruction can be replaced in its entirety. In other words, when the CPU recognizes that instruction 1 includes the modifiable instruction flag (i.e., when instruction 1 is in form 2), the CPU can replace instruction 1 with another instruction that includes the aforementioned modifiable instruction flag.

[0076] Optionally, the instruction identifier can be in any s+1 fields of instruction 1, and the field where the modifiable instruction flag is located can be in any j fields of instruction 1. The above Table 3 is only an example and does not constitute a limitation to this application.

[0077] When the form of instruction 1 is form 2, some or all CPUs (for example, CPU A and CPU B) in the multi-core system respectively store an identical instruction library, and each instruction in the instruction library includes the aforementioned modifiable instruction flag, that is, multiple instructions in the instruction library can be replaced with each other, and each instruction maps a set of context information.

[0078] For example, as shown in Table 4, the mapping relationship between instructions in the instruction library and context information includes:

[0079]

[0080] Table 4

[0081] In some examples, the instruction library may also pre-store at least one initial instruction, and there is no mapping relationship between the initial instruction and any identifier of the context information. For example, the instruction identifier of the initial instruction is A0, and the content of the ISA definition part of the initial instruction is x1.

[0082] Form 3:

[0083] As shown in Table 5, the kth field to the Rth field store mapping identifiers, which can be read and written by the CPU. The other fields are used to implement the contents defined by the ISA. For example, instruction 1 may include the following fields:

[0084]

[0085] Table 5

[0086] When the form of instruction 1 is form 3, any CPU can read and write the content in the mapping identifier in instruction 1.

[0087] Optionally, the field where the mapping identifier is located can be in any R-k+1 fields of instruction 1. The above Table 5 is only an example and does not constitute a limitation to this application.

[0088] When the form of instruction 1 is form 3, some or all CPUs (eg, CPU A and CPU B) in the multi-core system respectively store a mapping relationship table, which includes a correspondence between at least one mapping identifier and at least one set of context information.

[0089] For example, as shown in Table 6, the mapping relationship table may be:

[0090] Mapping ID Identification of context information 00 C1 01 C2 10 C3

[0091] Table 6

[0092] In some examples, an initial value (eg, 11) may be preset in the mapping identifier of the instruction, and there is no mapping relationship between the initial value and any identifier of the context information.

[0093] In one possible design, the coherent cache obtains and stores an ELF file, wherein the ELF file includes instruction 1.

[0094] In another possible design, such as Figure 4 As shown, the source file is compiled for the first time to obtain the original ELF file (stored in the consistency cache); the CPU (any CPU in the multi-core system) executes the ELF file; the CPU can also determine the location of the ELF file where instruction 1 can be embedded through some analysis methods; embed instruction 1 in the ELF file to obtain a modified ELF file; the CPU can store the modified ELF file (including instruction 1) in the consistency cache.

[0095] Optionally, the multi-core system further includes a memory, and the coherent cache can obtain the original ELF file through the memory.

[0096] It should be understood that the analysis method used by the CPU to determine the location of the pre-embedded instruction 1 in the ELF file can refer to conventional techniques in the art and is not limited by this application. For example, the CPU can determine the location of the pre-embedded instruction 1 in the ELF file based on hardware performance data (such as a performance monitor unit (PMU) call stack, BRBE, etc.) and profiling to locate modification points.

[0097] like Figure 4 As shown, the ELF file includes a ".text" section. The indexes of the instructions included in the ".text" section include program counters (PC) 1 to PC Y. The ".text" section can be used to store one or more pre-embedded modifiable instructions (instructions corresponding to the shaded area in the figure). For example, the instructions corresponding to PC 1 and PC Y-1 are all modifiable instructions.

[0098] Optionally, the ELF file may also include a "header" part, a "PLT" part, and other parts, which are not limited in this application.

[0099] S302: CPU A obtains the aforementioned instruction 1 through the coherent cache.

[0100] S303: CPU A obtains at least one set of context information stored in itself, wherein the context information in the at least one set of context information can be used to support CPUs other than CPU A to execute thread A.

[0101] For example, CPU A retrieves its stored context information to obtain the at least one set of context information. The specific process includes: Modification Module A can obtain context information stored in at least one micro-architecture module A and select, from all the obtained context information, context information that can be used for speculative execution of Thread A on other CPUs, thereby obtaining the at least one set of context information.

[0102] S304: CPU A obtains instruction 2 based on instruction 1.

[0103] Optionally, when CPU A detects that instruction 1 is in any of the three forms mentioned above, it executes the above-mentioned action of obtaining instruction 2 based on instruction 1.

[0104] Based on the three forms provided in S301 above, the process of CPU A executing S304 includes the following three possible ways:

[0105] Method 1:

[0106] CPU A may write at least one set of context information obtained in the aforementioned S303 into the reserved field of instruction 1 (instruction structure of form 1), thereby obtaining instruction 2.

[0107] Method 2:

[0108] Assume that instruction 1 is an initial instruction pre-stored in the instruction library, and the ISA definition portion of the initial instruction is x1. The at least one set of context information acquired by CPU A in S303 is context information B1. Based on the instruction library illustrated in Table 4, CPU A can replace instruction 1 with instruction 2. To ensure that other fields remain unchanged, instruction 2 can be instruction numbered A1.

[0109] Method 3:

[0110] Assuming that the mapping identifier in instruction 1 is a preset initial value, the identifier of the at least one set of context information obtained by CPU A in S303 is C3. Based on the mapping relationship table shown in Table 6, CPU A can modify the mapping identifier in instruction 1 to 10, thereby obtaining instruction 2.

[0111] In one possible example, when CPU A obtains new context information (different from the aforementioned at least one set of context information), CPU A may update instruction 2 based on the newly obtained context information. The method for updating instruction 2 may refer to the aforementioned method 1, method 2, or method 3.

[0112] Optionally, the scenarios in which CPU A obtains new context information include: after executing S304 and before executing S305, CPU A may generate new context information while executing the first thread; or, after executing S304 and before executing S305, CPU A supplements the context information stored in itself and retrieves new context information.

[0113] S305: During the migration of the first thread from CPU A to CPU B, CPU A sends the aforementioned instruction 2 to CPU B through the coherent cache. For example, CPU A stores instruction 2 in the coherent cache, and CPU B reads instruction 2 from the coherent cache.

[0114] S306: CPU B reads instruction 2, thereby obtaining the aforementioned at least one set of context information.

[0115] Based on the three forms provided in S301 above, the process of CPU B executing S306 includes the following three possible ways:

[0116] Method A:

[0117] Based on the writing method of the aforementioned method 1, CPU B can read the reserved field segment of instruction A, thereby obtaining the aforementioned set of at least one set of context information.

[0118] Method B:

[0119] Based on the instruction library shown in Table 4, CPU B can determine that the context information corresponding to instruction 2 (ie, the set of at least one set of context information) is context information B1 according to instruction 2 (ie, instruction numbered A1).

[0120] Method C:

[0121] Based on the mapping relationship table exemplified in Table 6 above, CPU B may read the mapping identifier 10 in instruction 2, thereby determining that the identifier of the aforementioned at least one set of context information is C3.

[0122] In one possible design, CPU B can determine whether the aforementioned at least one set of context information has been stored in CPU B; if it has been stored, no further processing is required; if it has not been stored yet, the modification module B in CPU B will store the aforementioned at least one set of context information in at least one micro-architecture module B respectively.

[0123] S307: CPU B performs speculative execution (eg, prefetching, branch prediction) based on the aforementioned at least one set of context information.

[0124] For example, Figure 5 The following is an example of a partial process flow during branch prediction for branch predictor A in CPU A and branch predictor B in CPU B. Both branch predictor A and branch predictor B include a program counter (PC) module, a global (local) history register (G(L)HR) module, and a decision maker. Optionally, branch predictor A and branch predictor B may also include a branch target buffer (BTB) module.

[0125] When a modifiable instruction enters the branch predictor A, the branch predictor A can index the historical information of the jump direction of the modifiable instruction in the history list (pattern history table, PHT) through PC and GHR or (xor, or XOR) query; the branch predictor A can also query its jump target based on the BTB to obtain a fixed format.

[0126] Furthermore, CPU A can compress (i.e., re-encode) the historical context information in the PHT and write it into the hint field of the modifiable instruction. Conversely, if the jump direction historical information cannot be found in the PHT, the branch prediction fails. Finally, when branch predictor B receives the re-encoded modifiable instruction, it can read the hint field of the modifiable instruction to obtain the history list corresponding to the PHT, that is, obtain the historical context information; this historical context information can be used to implement speculative execution.

[0127] In the various embodiments of the present application, unless otherwise specified or there is a logical conflict, the terms and / or descriptions between different embodiments are consistent and can be referenced by each other. The technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationships.

[0128] The method provided in the embodiment of the present application is described above in conjunction with the accompanying drawings. The thread migration device provided in the embodiment of the present application is described below in conjunction with the accompanying drawings.

[0129] Based on the same technical concept, the present application also provides a thread migration device, which is used to implement the thread migration method provided in the above embodiment. Figure 6 As shown, the thread migration device 600 includes a communication unit 601 and a processing unit 602. The communication unit 601 is used to receive and / or send data; the processing unit 602 is used to implement Figure 2 The steps in the thread migration method are shown.

[0130] In a possible example, the thread migration device 600 is used to implement the above Figure 2 When the function of the first processor is shown, the processing unit 602 is used to: generate a second instruction based on the first instruction and the first context information during the migration of the first thread from the first processor to the second processor; the first instruction is the instruction required to execute the first thread; the communication unit 601 is used to: send the second instruction to the second processor; the second instruction is used to support the second processor to execute the first thread.

[0131] In one possible design, the first instruction includes a modifiable field, and the processing unit 602 is specifically used to add the first context information to the modifiable field of the first instruction to obtain the second instruction.

[0132] In a possible design, the first instruction and the second instruction both include an instruction identifier and instruction content; the instruction identifiers of the first instruction and the second instruction are different, and the instruction contents of the first instruction and the second instruction are the same.

[0133] In one possible design, the processing unit 602 is specifically used to: obtain a first correspondence set between multiple pre-stored instructions and multiple context information; and select a second instruction from multiple instructions in the first correspondence set based on the first instruction and the first context information.

[0134] In one possible design, the processing unit 602 is specifically used to: determine at least one alternative instruction that has a corresponding relationship with the first context information among multiple instructions in the first correspondence set; and select an alternative instruction with the same instruction content as the first instruction as the second instruction among at least one alternative instruction.

[0135] In one possible design, the first instruction includes a modifiable mapping identifier field, and the processing unit 602 is specifically used to: obtain a second set of correspondences between multiple pre-stored mapping identifiers and multiple context information; in the second correspondence set, determine the target mapping identifier corresponding to the first context information; add the target mapping identifier to the mapping identifier field of the first instruction to obtain a second instruction.

[0136] In one possible design, the communication unit 601 is specifically configured to send the second instruction to the second processor through a consistent cache device.

[0137] In a possible design, the communication unit 601 is further configured to: obtain the first instruction through a consistent cache device.

[0138] In one possible design, the communication unit 601 is specifically used to: obtain second context information; the second context information is different from the first context information; the processing unit 601 is used to perform the following steps through the communication unit 602: update the second instruction according to the second context information, and send the updated second instruction to the second processor.

[0139] In one possible design, the first processor and the second processor are set in the same device.

[0140] In a possible example, the thread migration device 600 is used to implement the above Figure 2 When the function of the first processor is shown, the communication unit 601 is used to: receive a second instruction from the first processor during the process of the first thread migrating from the first processor to the second processor; the processing unit 602 is used to: determine the first context information corresponding to the first thread according to the second instruction; the first context information is used by the second processor to execute the first thread.

[0141] In one possible design, the communication unit 601 is specifically configured to receive the second instruction from the first processor through a coherent cache device.

[0142] Based on the same technical concept, the embodiment of the present application also provides another thread migration device 700, which can implement the thread migration method provided in the above embodiment. Figure 7 As shown, the thread migration device 700 includes a processor 701. Optionally, the thread migration device 700 also includes a memory 702 and / or a communication interface 703. The memory can be located inside the thread migration device or outside the thread migration device, which is not limited in this application. The communication interface 703, the processor 701, and the memory 702 are interconnected. Exemplarily, the thread migration device 700 can be the first processor or the second processor shown in the embodiment of this application.

[0143] Optionally, the communication interface 703, the processor 701, and the memory 702 are interconnected via a bus 704. The bus 704 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, and the like. For ease of representation, Figure 7 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0144] The communication interface 703 is used to receive and / or send signals to implement communication with other devices other than the thread migration apparatus.

[0145] The processor 701 may be used to execute the aforementioned Figure 2 Any of the thread migration methods in the embodiment can be referred to the description in the above embodiments and will not be described in detail here. The processor 701 can be a central processing unit (CPU), a network processor (NP), or a combination of a CPU and an NP, etc. The processor 701 can further include a hardware chip. The above-mentioned hardware chip can be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The above-mentioned PLD can be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof. When implementing the above-mentioned functions, the processor 701 can be implemented through hardware, and of course, the corresponding software implementation can also be executed by hardware.

[0146] The memory 702 is used to store program instructions, etc. Specifically, the program instructions may include program code, which includes computer operating instructions. The memory 702 may include random access memory (RAM) and may also include non-volatile memory (non-volatile memory), such as at least one disk storage device. The processor 701 executes the program instructions stored in the memory 702 to implement the above functions, thereby implementing the methods provided in the above embodiments.

[0147] Based on the same technical concept, an embodiment of the present application further provides a computer program, which, when executed on a computer, enables the computer to execute the method provided in the above embodiment.

[0148] Based on the same technical concept, an embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a computer, the computer executes the method provided in the above embodiment.

[0149] The storage medium may be any available medium that can be accessed by a computer. By way of example and not limitation, computer-readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer.

[0150] Based on the same technical concept, an embodiment of the present application further provides a chip, which is used to read a computer program stored in a memory to implement the method provided in the above embodiment.

[0151] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0152] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0153] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0154] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0155] Obviously, those skilled in the art may make various changes and modifications to the present application without departing from the scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is intended to include these modifications and variations.

Claims

1. A thread migration method, applied to a first processor, wherein the first processor stores first context information generated during the execution of a first thread, characterized in that: The method comprises: During migration of the first thread from the first processor to the second processor, generating a second instruction according to the first instruction and the first context information; the first instruction is an instruction required to execute the first thread; The second instruction is sent to the second processor; the second instruction is used to support the second processor in executing the first thread.

2. The method according to claim 1, wherein The first instruction includes a modifiable field, and generating the second instruction according to the first instruction and the first context information includes: The first context information is added to a modifiable field of the first instruction to obtain the second instruction.

3. The method according to claim 1, wherein The first instruction and the second instruction both include an instruction identifier and instruction content; the instruction identifiers of the first instruction and the second instruction are different, and the instruction contents of the first instruction and the second instruction are the same.

4. The method according to claim 3, wherein Generating a second instruction according to the first instruction and the first context information includes: Obtaining a first set of correspondences between a plurality of pre-stored instructions and a plurality of context information; The second instruction is selected from the multiple instructions in the first correspondence set according to the first instruction and the first context information.

5. The method according to claim 4, wherein The selecting, according to the first instruction and the first context information, the second instruction from the multiple instructions in the first correspondence set includes: Determining, among the multiple instructions in the first correspondence set, at least one candidate instruction that has a correspondence with the first context information; Among the at least one candidate instruction, an alternative instruction having the same instruction content as the first instruction is selected as the second instruction.

6. The method according to claim 1, wherein The first instruction includes a modifiable mapping identification field, and generating the second instruction according to the first instruction and the first context information includes: Obtaining a second set of correspondences between a plurality of pre-stored mapping identifiers and a plurality of context information; Determining a target mapping identifier corresponding to the first context information in the second correspondence set; The target mapping identifier is added to the mapping identifier field of the first instruction to obtain the second instruction.

7. The method according to any one of claims 1 to 6, wherein: The sending the second instruction to the second processor includes: The second instruction is sent to the second processor through the coherent cache device.

8. The method according to any one of claims 1 to 7, wherein: The method further comprises: The first instruction is obtained through a coherent cache device.

9. The method according to any one of claims 1 to 8, wherein: The sending the second instruction to the second processor includes: Acquire second context information; the second context information is different from the first context information; The second instruction is updated according to the second context information, and the updated second instruction is sent to the second processor.

10. The method according to any one of claims 1 to 9, wherein: The first processor and the second processor are set in the same device.

11. A thread migration method, applied to a second processor, characterized in that: The method comprises: receiving a second instruction from the first processor during migration of the first thread from the first processor to the second processor; First context information corresponding to the first thread is determined according to the second instruction; the first context information is used by the second processor to execute the first thread.

12. The method according to claim 11, wherein The receiving a second instruction from the first processor includes: A second instruction is received from the first processor through the coherent cache device.

13. A thread migration device, characterized in that: include: a communication unit and a processing unit; The communication unit is used to receive and / or send data; The processing unit is used to execute the method according to any one of claims 1 to 10, or to execute the method according to any one of claims 11 to 12.

14. A thread migration device, characterized in that: include: at least one processor; The at least one processor is configured to execute the method according to any one of claims 1 to 10, or to execute the method according to any one of claims 11 to 12.

15. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, which, when called by a computer, enable the computer to execute the method according to any one of claims 1 to 10, or to execute the method according to any one of claims 11 to 12.

16. A chip system, characterized in that: Including processor; The processor is used to execute a computer-executable program, so that a device equipped with the chip system is used to execute the method according to any one of claims 1 to 10, or to execute the method according to any one of claims 11 to 12.