Processor failure rate processing method and processing apparatus, and electronic device

By sharing address information and execution results between multiple processor cores, the problem of processor failure is solved, the overall performance and security performance of the processor are improved, and the area and cost of the chip are reduced.

WO2025130282A1PCT designated stage expired Publication Date: 2025-06-26HUAWEI TECH CO LTD

Patent Information

Application Number
PCT/CN2024/124626
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-22
Filing Date
2024-10-14
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

During use, the processor fails due to quality defects, environment, aging or improper use, etc., and a method is needed to solve the problem of CPU failure and improve the security performance of the CPU.

Method used

By sharing the address information obtained by high-performance processor cores to perform tasks among multiple processor cores, and performing services of other cores as prefetch information based on these address information, the accuracy of prefetch information is improved, thereby improving the processing performance of other cores. At the same time, the failure efficiency of the processor core is verified based on the execution results, and the security performance of the processor is improved.

Benefits of technology

This method not only improves the overall performance of the processor, but also verifies the failure efficiency of the processor core, enhances the security performance of the CPU and reduces the area and cost of the chip.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024124626_26062025_PF_FP_ABST
    Figure CN2024124626_26062025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of electronics, and provides a processor failure rate processing method and processing apparatus, and an electronic device, for use in improving the efficiency of processors. The processor failure rate processing method comprises: a first processor core sends address information to a second processor core, wherein the address information comprises a storage address of an instruction and data in a memory, the instruction and the data are used by the first processor core to execute a task, and the performance of the first processor core is higher than that of the second processor core; the second processor core acquires the instruction and the data from the memory by using the address information as prefetch information; and the second processor core executes the task on the basis of the instruction and the data to obtain an execution result, wherein the execution result is used for verifying the failure rate of at least one of the first processor core or the second processor core.
Need to check novelty before this filing date? Find Prior Art

Description

Processor failure rate processing method, processing device and electronic equipment

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on December 22, 2023, with application number 202311788375.4 and application name “A method for processing processor failure rate, a processing device and an electronic device”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of electronic technology, and in particular to a method for processing processor failure rate, a processing device, and an electronic device. Background Art

[0003] Artificial intelligence technology is widely used in fields such as smart cars, robotics, and aerospace. With the rapid development of AI technology, users are increasingly demanding higher levels of processor security and performance. A processor can include a central processing unit (CPU). The CPU serves as the computing and control core of a computer system and is the final execution unit for information processing and program execution. However, CPU failure rates, defined as CPU malfunctions, can occur during use due to quality defects, environmental factors, aging, or improper use. Therefore, a method to address this issue and improve CPU security is urgently needed.

[0004] Summary of the Invention

[0005] The present application provides a method for processing processor failure rate, a processing device, and an electronic device for improving the efficiency of the processor.

[0006] To achieve the above objectives, this application adopts the following technical solutions:

[0007] In a first aspect, a method for processing a processor failure rate is provided, the method comprising: a first processor core sends address information to a second processor core, the address information including the storage address in a memory of instructions and data used by the first processor core to execute a task, the performance of the first processor core being higher than that of the second processor core; the second processor core obtains instructions and data from the memory by using the address information as prefetch information (or speculative information); the second processor core executes the task according to the instructions and data to obtain an execution result, and the execution result is used to verify the failure rate of at least one of the first processor core or the second processor core.

[0008] In the technical solution provided by the present application, when the first processor core and the second processor core process the same business, the first processor core sends the storage address of the instructions and data required for executing the task to the second processor core, and the second processor core directly uses the storage address as prefetch information to obtain the instructions and data for executing the task from the memory. The second processor core can obtain the prefetch information for executing the task without performing a prefetch operation, thereby reducing the area, cost and power consumption of the device used to perform the prefetch operation in the second processor core, and at the same time reducing the area, cost and power consumption of the second processor core. When the second processor core is integrated into the chip, the area and cost of the chip can be reduced; on the other hand, based on the execution results of the two processor cores, the failure rate of at least one of the two processor cores can be verified, thereby improving the safety performance of the CPU. In addition, the address information is the storage address of the instructions and data required for executing the task. Using the address information as prefetch information (or speculative information) improves the accuracy of the prefetch information (speculative information). Further, the second processor core uses the instructions and data to execute the task, thereby improving the efficiency of executing the task, improving the processing performance of the second processor core, and thereby improving the overall performance of the second processor core.

[0009] In one possible implementation of the first aspect, the first processor core includes a cache queue; and the first processor core sending address information to the second processor core includes: the cache queue sending the address information to the second processor core. In the above possible implementation, using the cache queue to transmit the address information improves transmission efficiency.

[0010] In one possible implementation of the first aspect, the first processor core further includes a memory access control module; the method further includes: the memory access control module sending multiple storage addresses to a cache queue, where the multiple storage addresses are storage addresses accessed by the memory access control module during the execution of a task by the first processor core; and the cache queue filtering duplicate addresses from the multiple storage addresses to obtain address information. In this possible implementation, the cache queue filtering duplicate addresses from the multiple storage addresses improves the transmission efficiency of the address information and, further, reduces the storage capacity occupied by the cache address information in the second processor core when the address information is transmitted to the second processor core.

[0011] In a possible implementation of the first aspect, the first processor core sending address information to the second processor core includes: the first processor core sending the address information to the second processor core via a bus, where the bus is designed specifically for transmitting address information. In this possible implementation, using the specifically designed bus to transmit address information improves the rate at which the address information is transmitted.

[0012] In one possible implementation of the first aspect, the second processor core includes a cache; before the second processor core retrieves instructions and data from the memory using the address information as prefetch information, the method further includes: the second processor core writing the address information to the cache, and writing the instructions and data corresponding to the address information to the cache. In this possible implementation, the processor core retrieves instructions and data from the cache faster than from the memory, thereby increasing the instruction and data retrieval rate and improving the performance of the second processor core.

[0013] In a possible implementation of the first aspect, the method further includes: a verification module verifying a failure rate of at least one of the first processor core or the second processor core based on the execution result. In the above possible implementation, the problem of processor failure rate is solved and the safety performance of the processor is improved.

[0014] In a second aspect, a device for processing processor failure rate is provided, the processing device including a first processor core and a second processor core; the first processor core is used to send address information to the second processor core, the address information including the storage address of the instructions and data used by the first processor core to execute a task in a memory, the performance of the first processor core being higher than that of the second processor core; the second processor core is used to obtain instructions and data from the memory by using the address information as prefetch information, and execute the task according to the instructions and data to obtain an execution result, and the execution result is used to verify the failure rate of at least one of the first processor core or the second processor core.

[0015] In a possible implementation of the second aspect, the first processor core includes a cache queue, where the cache queue is configured to send address information to the second processor core.

[0016] In a possible implementation of the second aspect, the first processor core also includes a memory access control module; the memory access control module is used to send multiple storage addresses to the cache queue, and the multiple storage addresses are storage addresses accessed by the memory access control module during the execution of the task by the first processor core; the cache queue is used to filter out duplicate addresses in the multiple storage addresses to obtain address information.

[0017] In a possible implementation manner of the second aspect, the processing device further includes a bus; and the first processor core is configured to send address information to the second processor core through the bus.

[0018] In a possible implementation of the second aspect, the second processor core includes a cache; and the second processor core is further configured to: write address information into the cache before retrieving instructions and data from the memory by using the address information as prefetch information.

[0019] In a possible implementation manner of the second aspect, the processing device further includes: a verification module, configured to verify a failure rate of at least one of the first processor core or the second processor core according to the execution result.

[0020] In a third aspect, an electronic device is provided, which includes a memory and a processing device, the memory is used to store computer instructions, the processing device includes a first processor core and a second processor core, and the processing device is used to execute the method for processing processor failure rate provided by the first aspect or any possible implementation of the first aspect.

[0021] In a fourth aspect, a computer-readable storage medium is provided, which stores computer instructions. When the computer instructions are executed on a processing device, the processing device executes a method for processing processor failure rate as provided in the first aspect or any possible implementation of the first aspect.

[0022] In a fifth aspect, a computer program product comprising instructions is provided. When the computer program product is run on a processing device, the processing device is caused to execute a method for processing processor failure rate as provided in the first aspect or any possible implementation of the first aspect.

[0023] It can be understood that the processing device, electronic device, computer-readable storage medium and computer program product provided above can be used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method provided above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] FIG1 is a schematic structural diagram of a chip system;

[0025] FIG2 is a schematic diagram of the structure of another chip system;

[0026] FIG3 is a schematic structural diagram of another chip system;

[0027] FIG4 is a schematic diagram of the execution efficiency of a processor core provided by an embodiment of the present application;

[0028] FIG5 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application;

[0029] FIG6 is a schematic diagram of the structure of a processor provided in an embodiment of the present application;

[0030] FIG7 is a flowchart of a method for processing processor failure rate provided by an embodiment of the present application;

[0031] FIG8 is a flowchart of another method for processing processor failure rate provided by an embodiment of the present application;

[0032] FIG9 is a schematic structural diagram of a device for processing processor failure rate provided in an embodiment of the present application. DETAILED DESCRIPTION

[0033] In this application, "at least one" refers to one or more, and "more" refers to two or more. "And / or" describes the relationship between associated objects, indicating that three relationships can exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the associated objects are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, "at least one of a, b, or c" can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple. In addition, the embodiments of this application use words such as "first" and "second" to distinguish between identical or similar items with substantially the same function and effect. For example, the first threshold and the second threshold are merely to distinguish different thresholds and do not define their order of precedence. Those skilled in the art will understand that words such as "first" and "second" do not define the quantity or execution order.

[0034] It should be noted that, in this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described in this application as "exemplary" or "for example" should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0035] Before introducing the embodiments of the present application, the relevant knowledge of dual-core lockstep is first introduced.

[0036] Artificial intelligence technology is widely used in fields such as smart cars, robotics, and aerospace. With the rapid development of AI technology, users are increasingly demanding higher levels of processor security and performance. Processors include central processing units (CPUs). The CPU serves as the computing and control core of a computer system and is the final execution unit for information processing and program execution. However, CPU failure rates, often due to quality defects, environmental factors, aging, or improper use, can lead to CPU malfunctions. Therefore, there is an urgent need for methods to address this CPU failure rate and improve CPU performance.

[0037] Currently, dual core lock step (DCLS) can be used to address CPU failure rates and improve CPU performance. DCLS can be implemented using the following three methods:

[0038] Solution 1, as shown in FIG1 , is a schematic diagram of the structure of a processor 10. Processor 10 includes a CPU core 110, a CPU core 120, an input / output port 130, and a verification module 140. Furthermore, cache 150 and main memory 160 corresponding to CPU core 110, and cache 170 and main memory 180 corresponding to CPU core 120, are provided. I / O port 130 provides an interface between processor 10 and peripheral interface modules, such as a keyboard, mouse, or Universal Serial Bus (USB) device. Processor 10 is configured to write instructions and data received by I / O port 130 into main memory. For example, the instructions and data may have been pre-written by a user via a keyboard. Main memory may be DDR (double data rate) memory. Before executing a task, each CPU core reads the instructions and data corresponding to the task from the corresponding main memory and writes the read instructions and data into the corresponding cache, such as an on-chip cache located on the same chip as the CPU. Each CPU core executes a task based on the instructions and data in the corresponding cache, and writes the execution results to the corresponding main memory through a write request. The write request includes the execution result and the address, which is the address corresponding to the storage space in the main memory used to store the execution result. At the same time, the CPU core sends a write request to the verification module 140. The verification module 140 (e.g., a parallel comparator) compares the two most recently written execution results in the memory every clock cycle (e.g., 60s) to see if they are consistent. For example, the verification module 140 uses error correction code (ECC) or error detection and correction (EDC) to compare whether the execution results are consistent. If the execution results of CPU core 110 and CPU core 120 are inconsistent, it means that one of the CPU cores has failed. The fault handling module is used to handle the fault, thereby solving the problem of CPU failure rate. In the example shown in Figure 1, dual-core lockstep is implemented by CPU core 110, CPU core 120, cache 150, cache 170, main memory 160, and main memory 180. Figure 1 only shows a partial structure of the processor.

[0039] Solution 2, as shown in FIG2, FIG2 is a schematic diagram of the structure of another processor 20, which includes a CPU core 210, a CPU core 220, an input / output port 230, a verification module 240, a main memory 250, a cache 260 corresponding to the CPU core 210, and a cache 270 corresponding to the CPU core 220. Among them, the main memory 250 is shared by the CPU core 210 and the CPU core 220. The example shown in FIG2 is a dual-core lockstep implemented by the CPU core 210, the CPU core 220, the cache 260, and the cache 270. The specific process of implementing the dual-core lockstep is similar to the process of implementing the dual-core lockstep by the processor 10 shown in FIG1, and will not be repeated here.

[0040] Solution 3, as shown in FIG3, FIG3 is a schematic diagram of the structure of another processor 30, which includes a CPU core 310, a CPU core 320, an input / output port 330, a verification module 340, a main memory 350, and a cache 360. Among them, the main memory 350 and the cache 360 ​​are shared by the CPU core 210 and the CPU core 220. The example shown in FIG3 is a dual-core lockstep implemented by the CPU core 310 and the CPU core 320. The specific process of implementing the dual-core lockstep is similar to the process of implementing the dual-core lockstep by the processor 10 shown in FIG1, and will not be repeated here.

[0041] Optionally, some components of the processor shown in Figures 1, 2 and 3 (for example, CPU core and cache, etc.) can be integrated on the chip, and the processor shown in Figures 1, 2 and 3 can include more CPU cores. For example, the processor shown in Figure 1 can include 3 CPU cores. Figures 1, 2 and 3 are only examples and do not constitute a limitation.

[0042] However, in the above scheme, in order to achieve dual-core lockstep, it is necessary to set up at least two identical CPU cores, and at least two identical CPU cores perform the same task. The results of different CPU cores performing the same task (for example, message transmission, message modification, image denoising tasks or picture processing tasks, etc.) are used to determine whether the CPU core is faulty. However, setting up dual-core lockstep (that is, setting up two identical CPU cores in one chip) increases the area and cost of the chip system.

[0043] Based on this, an embodiment of the present application provides a method for processing processor failure rate. When multiple processor cores process the same business (or so-called homogeneous business), this method can share the address information obtained by the high-performance processor core executing the task between the cores, and execute the business of other cores based on the address information (which can be used as pre-fetch information or speculative information), thereby improving the accuracy of the pre-fetch information and thus improving the processing performance of other cores; based on the execution results, the implementation efficiency of at least one processor core among the multiple processor cores is verified to improve the safety performance of the processor, thereby improving the overall performance of the processor.

[0044] For example, as shown in FIG4 , the arrows in FIG4 represent processor cores, and the length of each arrow represents the execution efficiency of the processor core. In the first method, multiple processor cores with the same performance (for example, the CPU cores shown in FIG1 , FIG2 , and FIG3 ) are used to perform the same task, and the execution efficiency of the multiple processor cores is the same; in the second method, a first processor core with higher performance and other processor cores with lower performance than the first processor core are used to perform the same task. In the first stage, the execution efficiency of the first processor core is faster than that of the other processor cores. At this time, the first processor core shares address information with the other cores; in the second stage, after the other cores execute the business according to the shared address information, the execution efficiency finally achieved is close to that of the first processor core, thereby improving the processing performance of the processor.

[0045] The technical solution provided in the embodiment of the present application can be applied to electronic devices including high-security processors, and the electronic devices may include but are not limited to mobile robots, drones, vehicle-mounted equipment, aerospace equipment, etc. The structure of the electronic device is described below in conjunction with Figure 5. For example, Figure 5 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device 50 may include: a processor 510, a bus 520, a memory 530, and a communication interface 540. The processor 510, the memory 530, and the communication interface 540 are connected via a bus 520.

[0046] It should be understood that in this embodiment, the processor 510 is the control center of the electronic device 50. It uses various interfaces and buses 520 to connect the various parts of the entire device. By running or executing software programs and / or software modules stored in the memory 530 and calling data stored in the memory 530, it performs various functions of the electronic device 50 and processes data, thereby controlling the electronic device 50 as a whole. The processor 510 can be a CPU, and the processor 510 can also be other general-purpose processors, digital signal processors (digital signal processing, DSP), (application-specific integrated circuit, ASIC), field-programmable gate arrays (field-programmable gate array, FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc. The processor 510 can also be a graphics processing unit (GPU), a neural network processing unit (NPU), a microprocessor, an ASIC, or one or more integrated circuits used to control the execution of the program of the present application solution.

[0047] In an embodiment of the present application, the processor 510 may be a multi-core processor. For example, FIG6 is a schematic diagram of the structure of a processor 510 provided in an embodiment of the present application. As shown in FIG6 , the processor 510 may include multiple processor cores (also referred to as kernels). For example, the processor 510 may include a first processor core and a second processor core. The first processor core includes a memory access control module and a cache queue. The memory access control module may also be referred to as a load store unit (LSU), and the cache queue may be a prefetch input buffer (inbuf); further, the first processor core may also include a level 1 cache (L1), a level 2 cache (L2), a hardware prefetcher (HWP), a prefetch to level 1 cache (PFL1) interface targeting the first cache, a prefetch to level 2 cache (PFL2) interface targeting the second cache, and a prefetch to level 3 cache (PFL3) interface targeting the third cache, and other modules. The figure uses the example of a first processor core including an LSU, inbuf, L1, L2, HWP, PFL1 interface, PFL2 interface, and PFL3 interface, and the LSU and L1 are represented as LSU+L1. Optionally, as shown in FIG6 , other processor cores (slave cores), such as the second processor core, can have a similar structure to the first processor core. FIG6 illustrates the example of processor 510 including the first and second processor cores, and does not limit the number of processor cores.

[0048] Optionally, the processor 510 may include a bus and a verification module. The bus may be used to transmit address information, wherein the bus may be a multiplexed bus or a newly added bus. The verification module may be used to compare whether the execution result (output result) of the first processor core is consistent with the execution result of the second processor core. For example, the verification module may be a parallel comparator. Figure 6 only illustrates a partial structure of the processor 510. Those skilled in the art will understand that the structure of the processor shown in Figure 6 does not constitute a limitation on the processor, and may include more or fewer components than shown, or combine certain components, or arrange the components differently.

[0049] In a possible embodiment, the processor 510 shown in FIG. 6 may be applied in an electronic device in the form of a chip. For example, the processor core and storage included in the processor 510 may be provided on the chip.

[0050] 5 , the communication interface 540 is used to implement communication between the electronic device 50 and an external device or component.

[0051] Bus 520 may include a path for transmitting information between the aforementioned components (e.g., processor 510 and memory 530). Bus 520 may include an address bus, a data bus, a control bus, etc. However, for clarity, various buses are labeled as bus 520 in the figure. Bus 520 may be a Peripheral Component Interconnect Express (PCIe) bus, an Extended Industry Standard Architecture (EISA) bus, a unified bus (Ubus or UB), a Compute Express Link (CXL), a Cache Coherent Interconnect for Accelerators (CCIX), etc.

[0052] It is worth noting that Figure 5 only takes the electronic device 50 including one processor (for example, processor 510) and one memory (for example, memory 530) as an example. Here, the processor 510 and the memory 530 are respectively used to indicate a type of device or equipment. In a specific embodiment, the number of each type of device or equipment can be determined according to business requirements.

[0053] The memory 530 can be used to store data, software programs, and software modules. It mainly includes a program storage area and a data storage area. The program storage area can store an operating system and at least one application required for a function, such as a sound playback function or an image playback function. The data storage area can store data created based on the use of the electronic device, such as audio data, image data, or table data. For example, the memory 530 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0054] Although not shown, the electronic device may further include an audio component and a communication component, for example, the audio component includes a microphone, and the communication component includes a wireless fidelity (WiFi) module or a Bluetooth module, etc., which will not be described in detail in the embodiments of the present application. It will be understood by those skilled in the art that the electronic device structure shown in FIG5 does not constitute a limitation on the electronic device, and may include more or fewer components than shown, or combine certain components, or arrange the components differently.

[0055] Before introducing the method for processing the processor failure rate provided by the embodiment of the present application, the application scenario of the method for processing the processor failure rate provided by the embodiment of the present application is first introduced and explained. The method for processing the processor failure rate provided by the embodiment of the present application is applied to the processor 510 shown in Figure 5 above. In a scenario with high security level requirements, the first processor core and the second processor core in the processor 510 can be used to execute the same task at the same time. Since the performance of the first processor core is higher / stronger than that of the second processor core, that is, the execution efficiency of the first processor core in executing the task is higher than the execution efficiency of the second processor core in executing the task, the first processor core can send the storage address of the instruction and data used to execute the task in the memory to the second processor core, and the second processor core can use the storage address as the pre-fetch information (or speculative information) of the task to execute the task, and send a request carrying the execution result to the verification module, which is used to instruct the execution result to be written to the memory 530. The verification module verifies whether the execution result of the first processor core and the execution result of the second processor core are consistent, so as to determine the failure rate of at least one of the first processor core and the second processor core, and uses the fault handling module to perform fault handling to solve the problem of processor failure rate, so as to improve the safety performance of the processor.

[0056] The following describes the method for processing processor failure rates provided in an embodiment of the present application, in conjunction with the processor shown in FIG6 . The processor may include multiple processor cores, and a multi-core processor may include a first processor core and a second processor core. The processor may include a CPU, a GPU, or an NPU, etc., which is not specifically limited in the embodiment of the present application.

[0057] FIG7 is a flow chart of a method for processing processor failure rate according to an embodiment of the present application. The method includes:

[0058] S701: The first processor core sends address information to the second processor core. The address information includes instructions used by the first processor core to execute tasks and storage addresses of data in a memory. The performance of the first processor core is higher than that of the second processor core.

[0059] The first processor core is a processor core with higher performance among the multiple processor cores, that is, the first processor core is a processor core with higher execution efficiency among the multiple processor cores. The first processor core can also be called a master core or a leading core. The second processor core is any processor core among the multiple processor cores whose performance is inferior to that of the first processor core, that is, the second processor core is any processor core among the multiple processor cores whose execution efficiency is inferior to that of the first processor core. The second processor core can also be called a slave core or a follower core. The memory can be the memory 530 shown in Figure 5 above. In actual applications, the first processor core can be a large core including many hardware resources, and the second processor core can be a small core including few hardware resources, that is, a simplified small core.

[0060] Optionally, the first processor core may be pre-set or designated as the leader core. Exemplarily, the leader core is determined based on the physical location of the multiple processor cores, such as by setting the first core in a physical location as the leader core. Alternatively, the leader core is designated by software running on the processor. This embodiment of the present application does not specifically limit the method for designating the leader core.

[0061] Secondly, the task is the same task executed by the first processor core and the second processor core. The same task means that the data and / or instructions that the two tasks need to access during the execution are the same, for example, the instructions, data or branches to be executed are the same. In actual applications, the first processor core and the second processor core are used to execute the same task at the same time, that is, the first processor core and the second processor core adopt dual-core lockstep to solve the problem of processor failure rate. The first processor core can send the address information actually used to execute the task to the second processor core after determining the address information. Specifically, the address information can be sent during the execution of the task or after the task is completed. This embodiment does not limit this.

[0062] In a possible embodiment, when the first processor core receives an instruction to execute a task, the first processor core executes the task and accesses the memory or cache (for example, the first level cache or the second level cache) through the memory access control module to obtain the instructions and data required to execute the task and obtain address information; similarly, when the second processor core receives an execution instruction for a service, the second processor core executes the service, wherein the running speed of the first processor core is greater than the running speed of the second processor core, that is, the running speed of the first processor core is faster than the running progress of the second processor core, so that the first processor core sends address information to the second processor core, which can improve the speculation performance of the second processor core and thereby improve the performance of the second processor core in processing tasks.

[0063] In addition, the address information may include instruction address information and data address information. The instruction address information is the storage address in the memory of the instruction required by the first processor core to perform the task. The instructions required by the first processor core to perform the task may include, but are not limited to, front-end instruction fetch instructions, jump instructions (including jump direction and jump destination), back-end memory access instructions, sequential instructions, and value prediction instructions. The data address information is the storage address of the data required by the first processor core to perform the task.

[0064] In one possible embodiment, before the first processor core executes a task, a hardware prefetcher in the first processor core can be used to pre-write the instructions and data required for executing the task into the cache memory corresponding to the first processor core (e.g., L1 and L2) in units of cache lines. A cache line includes instructions and data, as well as the storage addresses corresponding to the instructions and data in memory. A cache line is typically 512 bits. While the first processor core is executing a task, the memory access control module sends an access request to the memory. The access request includes a storage address in the memory. The access request is typically 8 bits. The memory access control module (e.g., LUS) indirectly accesses the memory through the cache. Each time the memory is accessed, it needs to traverse the cache lines in the cache to check whether the storage address to be accessed is in the cache line. If the storage address is in the cache line and the instructions and data corresponding to the storage address are valid, the instructions and data stored in the cache line are directly read. If the storage address is not in the cache line, or if the storage address is in the cache line but the instructions or data corresponding to the storage address are invalid, the instructions and data are directly read from the memory into the cache and then read from the cache. During this process, a cache line can be accessed multiple times by the memory access control module, that is, a storage address can be accessed multiple times.

[0065] In one possible example, after the first processor core determines multiple storage addresses, the memory access control module (e.g., LUS) sends the multiple storage addresses to the cache queue (e.g., inbuf), and the multiple storage addresses are storage addresses accessed by the memory access control module during the execution of the task by the first processor core. The cache queue stores and filters duplicate addresses in the multiple storage addresses to obtain address information, and the cache queue sends the address information to the second processor core. For example, the cache queue sends the address information through the PFL3 interface. The cache queue can be a multi-input and one-output cache queue. In addition to filtering duplicate addresses, it can also be used to balance the bandwidth difference between input and output to improve the transmission efficiency of address information. Therefore, as shown in Figure 8, the method provided in the embodiment of the present application also includes:

[0066] S01: The memory access control module sends multiple storage addresses to the cache queue. The multiple storage addresses are storage addresses accessed by the memory access control module during the execution of a task by the first processor core. The cache queue filters duplicate addresses in the multiple storage addresses to obtain address information.

[0067] Step S701 thus includes: the cache queue sends address information to the second processor core.

[0068] For example, the address information may be a memory access request address that does not hit the first-level cache, or a subset or full set of first-level high-speed memory access addresses of any rule, etc., and the embodiments of the present application do not impose specific restrictions on this.

[0069] In one possible embodiment, the second processor core receives address information from the first processor core. Optionally, the multiple processor cores are coupled via a bus, which can be an existing bus in the processor (i.e., a bus in a multiplexed processor) or a bus separately provided by the present application (e.g., a newly added bus). The present embodiment does not impose any specific restrictions on this. Therefore, in the method provided by the embodiment of the present application, step S701 specifically includes: the first processor core sends address information to the second processor core via the bus. In this embodiment, the address information is transmitted via a separately provided bus, thereby improving the transmission efficiency of the address information.

[0070] Optionally, as shown in FIG6 , the structure of the second processor core is similar to that of the first processor core. The second processor core may include a memory access control module and a cache queue, which may be a prefetch input cache (inbuf). Furthermore, the second processor core may include other modules, such as a Level 1 cache (L1), a Level 2 cache (L2), a hardware prefetcher (HWP), a PFL1 interface, a PFL2 interface, and a PFL3 interface. Furthermore, the PFL3 interface of the second processor core is coupled to the PFL3 interface of the first processor core via a bus.

[0071] In a possible example, as shown in FIG6 , the second processor core receives address information from the first processor core via a bus. Specifically, a hardware prefetcher (eg, HWP) in the second processor core receives address information from the first processor core via a bus.

[0072] It can be understood that the multiple processor cores may include multiple slave cores, that is, in addition to the first processor core, the multiple processor cores may also include other slave cores similar to the first processor core. The first processor core (that is, the leader core) can send address information to each slave core. Figure 6 is used as an example to illustrate that the multiple slave cores include one slave core.

[0073] Furthermore, the method provided in the embodiment of the present application also includes:

[0074] S02: The second processor core writes the address information into a cache corresponding to the second processor core.

[0075] S702: The second processor core obtains the instruction and the data from the memory by using the address information as prefetch information.

[0076] In one possible embodiment, the second processor core may include a cache and a hardware prefetcher (e.g., HWP). The hardware prefetcher may be configured to receive address information, convert the received address information into a physical address corresponding to a memory, and use the physical address as prefetch information to retrieve instructions and data corresponding to the task from the memory. The hardware prefetcher may also be configured to write the physical address, as well as the retrieved instructions and data, into a cache (e.g., L1) corresponding to the second processor core.

[0077] Optionally, since the address information may include instruction address information and data address information, the corresponding physical address obtained by converting the address information may also include an instruction physical address and a data physical address. The instruction physical address is the storage address of the instruction corresponding to the task, and the data physical address is the storage address of the data corresponding to the task. The instruction physical address and the data physical address may be stored in different caches, respectively. For example, the cache corresponding to the second processor core may include a first cache and a second cache. The first cache may be used to store instructions, and the second cache may be used to store data. The instruction physical address may be stored in the first cache. Optionally, the first cache is also used to store instructions corresponding to the task; the data physical address may be stored in the second cache. Optionally, the second cache may be used to store data corresponding to the task.

[0078] Among them, the cache can be any one of the first-level cache (L1), second-level cache (L2) or third-level cache (L3) corresponding to the second processor core. All data or instructions stored in each level of cache are part of the next level of cache. The cache closer to the second processor core is faster and smaller. For example, the first-level cache is close to the second processor core, and the L1 cache is the cache with the smallest capacity and the fastest read and write speed among the three levels of cache. The second processor core can write the address information into the first-level cache, the second-level cache or the third-level cache, and the embodiments of the present application do not make specific limitations on this. The processor shown in Figure 6 takes the second processor core including L1 and L2 as an example.

[0079] Furthermore, the method provided in the embodiment of the present application also includes:

[0080] S703: The second processor core executes the task according to the instruction and the data to obtain an execution result, which is used to verify the failure rate of at least one of the first processor core or the second processor core.

[0081] Specifically, during the execution of a task, the second processor core directly obtains the instructions and data used to execute the task from the cache corresponding to the second processor core as pre-fetch information, and executes the task according to the obtained instructions and data to obtain the execution result. The second processor core sends the execution result to the verification module (for example, a parallel comparator).

[0082] Furthermore, the method provided in the embodiment of the present application further includes:

[0083] S704: The verification module verifies the failure rate of at least one of the first processor core or the second processor core according to the execution result.

[0084] Furthermore, the second processor core writes the execution result into the memory through a first write request, the first write request includes the execution result of the second processor core and a first address, the first address being the address corresponding to the storage space in the memory used to store the execution result of the second processor core. The second processor core simultaneously sends the first write request to the verification module. Similarly, the first processor core writes the execution result into the memory through a second write request, the second write request includes the execution result of the first processor core and a second address, the second address being the address corresponding to the storage space in the memory used to store the execution result of the first processor core, and the first processor core simultaneously sends the second write request to the verification module. The first address and the second address can be the same address, that is, the first address and the second address can be used to indicate the same storage space in the memory.

[0085] The verification module compares the two most recently written execution results in the memory at each clock cycle (for example, 60 seconds) to see if they are consistent. For example, the verification module can use error correction code (ECC) or error detection and correction (EDC) to verify whether the execution results (including data, address, and control lines) are consistent. If the execution results are consistent, the first processor core and the second processor core are normal; if the execution results are inconsistent, at least one of the first processor core and the second processor core is abnormal (including the first processor core being abnormal, the second processor core being abnormal, or both the first processor core and the second processor core being abnormal), a fault signal is issued, and the fault handling module is used to handle the fault, thereby solving the problem of processor failure rate and improving the safety performance of the processor.

[0086] In an embodiment of the present application, when the first processor core and the second processor core are processing the same business, the first processor core sends address information to the second processor core, and the second processor core uses the address information directly as prefetch information (or speculative information) to obtain instructions and data for executing the task from the memory. The second processor core can obtain prefetch information without performing a prefetch operation, which reduces the area overhead of the hardware prefetcher and, on the other hand, improves the accuracy of the prefetch information, thereby improving the processing performance of the second processor. In addition, the execution result is used to verify the failure rate of at least one of the two processor cores, that is, to detect CPU failure. This improves the safety performance of the processor core.

[0087] Furthermore, the above embodiment is described by taking the operating speed of the first processor core as higher than the operating speed of the second processor core as an example. In fact, this setting is not used to limit the application scenario of the embodiment. In actual applications, multiple processor cores can have the same or different capabilities or processing speeds, and which one has stronger capabilities is not limited, as long as the operating information obtained after one of the processor cores executes the task can be used to share with the other processor core in order to execute the technical solution mentioned in this embodiment. For example, still taking Figure 6 as an example, the functions of the leader core and the slave core can be swapped, and the second processor core (slave core) can send its operating information to the first processor core (leader core) so that the first processor core executes a similar embodiment process as described in Figure 7 to achieve similar functions. For example, the second processor core may first execute at least part of the function of the task and obtain address information due to software scheduling or user selection. The address information can be shared with the first processor core for continuing to execute the same task in order to achieve a similar effect. It can be understood that the leader core and slave core mentioned in this embodiment, and the difference between the two cores, are only a scenario that can be applied, but are not used to limit the technical solution.

[0088] Based on this, an embodiment of the present application further provides a device for processing processor failure rate, as shown in FIG9 , which can be applied to a processor. As shown in FIG9 , the device includes a first processor core and a second processor core. In an embodiment of the present application, the first processor core can be used to execute steps S701 and S01 in the above method embodiment, and / or other steps described herein; the second processor core can be used to execute steps S02, S702, and S703 in the above method embodiment, and / or other steps described herein.

[0089] The verification module verifies a failure rate of at least one of the first processor core or the second processor core according to the execution result.

[0090] It can be understood that the specific structure of the first processor core and the second processor core, as well as all relevant contents of the steps involved in the above method embodiment can be referred to the embodiment of the business processing device, and the embodiment of this application will not be repeated here.

[0091] In another aspect of the present application, an electronic device is provided, comprising a memory and at least one processor, the memory being configured to store computer instructions, and the at least one processor comprising multiple processor cores configured to execute the computer instructions, so as to enable the electronic device to implement the method for processing processor failure rates as provided above. Optionally, the at least one processor includes the apparatus for processing processor failure rates as provided above.

[0092] It can be understood that all relevant contents of each step involved in the above method embodiment can be referred to the embodiment of the processing device for processor failure rate and the embodiment of the electronic device, and the embodiment of the present application will not be repeated here.

[0093] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules or units is merely a logical functional division. In actual implementation, other division methods may be used, such as combining or integrating multiple units or components into another device, or ignoring or not implementing certain features.

[0094] The units described as separate components may or may not be physically separate, and the components shown as units may be one physical unit or multiple physical units, that is, they may be located in one place or distributed in multiple places. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0095] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. The readable storage medium may include: a USB flash drive, a mobile hard drive, a read-only memory, a random access memory, a magnetic disk, or an optical disk, etc., which can store program code. Based on this understanding, the technical solution of the embodiment of the present application, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product.

[0096] In another embodiment of the present application, a readable storage medium is also provided, which stores computer execution instructions. When a device (which may be a single-chip microcomputer, chip, etc.) or a processor executes the steps in the above method embodiment.

[0097] In another embodiment of the present application, a computer program product is provided, which includes computer instructions stored in a readable storage medium; at least one processor of the device can read the computer instructions from the readable storage medium, and at least one processor executes the computer instructions so that the device performs the steps in the above method embodiment.

[0098] Finally, it should be noted that the above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A method for processing processor failure rate, characterized in that: The method comprises: The first processor core sends address information to the second processor core, the address information including instructions used by the first processor core to execute tasks and storage addresses of data in a memory, and the performance of the first processor core is higher than that of the second processor core; The second processor core obtains the instruction and the data from the memory by using the address information as pre-fetch information; The second processor core executes the task according to the instruction and the data to obtain an execution result, and the execution result is used to verify the failure rate of at least one of the first processor core or the second processor core.

2. The method according to claim 1, characterized in that The first processor core includes a cache queue; The first processor core sends address information to the second processor core, including: The cache queue sends the address information to the second processor core.

3. The method according to claim 2, characterized in that The first processor core also includes a memory access control module; The method further comprises: The memory access control module sends a plurality of storage addresses to the cache queue, wherein the plurality of storage addresses are storage addresses accessed by the memory access control module during the process of the first processor core executing the task; The cache queue filters duplicate addresses among the multiple storage addresses to obtain the address information.

4. The method according to any one of claims 1 to 3, characterized in that: The first processor core sends address information to the second processor core, including: The first processor core sends the address information to the second processor core through a bus.

5. The method according to any one of claims 1 to 4, characterized in that: The second processor core includes a cache; Before the second processor core obtains the instruction and the data from the memory by using the address information as pre-fetch information, the method further includes: The second processor core writes the address information into the cache.

6. The method according to any one of claims 1 to 5, characterized in that: The method further comprises: The verification module verifies a failure rate of at least one of the first processor core or the second processor core according to the execution result.

7. A device for processing processor failure rate, characterized in that: The processing device includes a first processor core and a second processor core; The first processor core is used to send address information to the second processor core, the address information including instructions used by the first processor core to execute tasks and storage addresses of data in a memory, and the performance of the first processor core is higher than that of the second processor core; The second processor core is used to obtain the instruction and the data from the memory by using the address information as pre-fetch information, and execute the task according to the instruction and the data to obtain an execution result, and the execution result is used to verify the failure rate of at least one of the first processor core or the second processor core.

8. The processing device according to claim 7, characterized in that The first processor core includes a cache queue, The cache queue is used to send the address information to the second processor core.

9. The processing device according to claim 8, characterized in that The first processor core also includes a memory access control module; The memory access control module is used to send a plurality of storage addresses to the cache queue, wherein the plurality of storage addresses are storage addresses accessed by the memory access control module during the process of the first processor core executing the task; The cache queue is used to filter duplicate addresses among the multiple storage addresses to obtain the address information.

10. The processing device according to any one of claims 7 to 9, characterized in that: The processing device also includes a bus; The first processor core is used to send the address information to the second processor core through the bus.

11. The processing device according to any one of claims 7 to 10, characterized in that: The second processor core includes a cache; The second processor core is further configured to: write the address information into the cache before acquiring the instruction and the data from the memory by using the address information as pre-fetch information.

12. The processing device according to any one of claims 7 to 11, characterized in that: The processing device also includes: A verification module is used to verify the failure rate of at least one of the first processor core or the second processor core according to the execution result.

13. An electronic device, characterized in that: The electronic device includes a memory and a processing device, the memory is used to store computer instructions, the processing device includes a first processor core and a second processor core, and the processing device is used to execute the computer instructions to implement the processor failure rate processing method as described in any one of claims 1-6.

14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and when the computer instructions are executed on a processing device, the processing device is caused to execute the method according to any one of claims 1 to 6.

15. A computer program product comprising instructions, characterized in that When the computer program product is executed on a processing device, the processing device is caused to execute the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Processing method and processing device for failure rate of processor and electronic equipment

    CN120196468A

  • Full-hardware dual-core lock step processor fault tolerance system

    CN111581003A

  • Microprocessor architecture and microprocessor fault detection method

    CN114416435A

  • Cache consistency verification method and device, equipment and medium

    CN116775133A

  • Lock step control device and method for processor

    CN116821038A

Cited By

  • Calculation scheduling method and device for AI chip, electronic equipment and storage medium

    CN121277566A