Processing method and processing device for failure rate of processor and electronic equipment
By setting the first processor core and the second processor core in the processor, using address information as prefetch information to obtain instructions and data from the memory, the problem of processor failure is solved, and the security performance and overall performance of the processor are improved.
Patent Information
- Application Number
- CN202311788375.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-22
- Publication Date
- 2025-06-24
AI Technical Summary
During use, the processor fails due to quality defects, environment, aging or improper use, resulting in CPU failure, and the existing technology is difficult to effectively solve this problem.
By setting the first processor core and the second processor core in the processor, the first processor core sends the storage addresses of instructions and data for performing tasks to the second processor core, which obtains the instructions and data from the memory as prefetch information and performs the tasks to verify the failure efficiency of the processor core.
This method improves the security performance of the processor, reduces the area, cost, and power consumption of the devices used to perform prefetch operations in the second processor core, and improves the overall performance of the processor.
Smart Images

Figure CN120196468A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of electronic technologies, and in particular, to a method and apparatus for processing the failure rate of a processor and an electronic device. Background Art
[0002] Artificial intelligence technology has been widely applied in fields such as intelligent vehicles, robots, and aerospace. With the rapid development of artificial intelligence technology, users have higher and higher requirements for the security level and performance of processors. A processor may include a central processing unit (CPU). As the operation and control core of a computer system, the CPU is the final execution unit for information processing and program running. However, during use, due to quality defects, environment, aging, or improper use, etc., the failure rate of the CPU may occur. The failure rate of the CPU refers to CPU failures. Therefore, there is an urgent need for a method to solve the problem of the CPU failure rate and improve the security performance of the CPU. Summary of the Invention
[0003] This application provides a method and apparatus for processing the failure rate of a processor and an electronic device, which are used to improve the efficiency of the processor.
[0004] To achieve the above object, this application adopts the following technical solutions:
[0005] In a first aspect, a method for processing the failure rate of a processor is provided. The method includes: a first processor core sends address information to a second processor core, where the address information includes the storage addresses of instructions and data used by the first processor core to execute tasks in a memory, and the performance of the first processor core is higher than that of the second processor core; the second processor core obtains instructions and data from the memory by using the address information as prefetch information (or speculative information); the second processor core executes a task according to the instructions and data to obtain an execution result, and the execution result is used to verify the failure rate of at least one of the first processor core and the second processor core.
[0006] In the technical solution provided by this application, when the first processor core and the second processor core process the same service, the first processor core sends the storage addresses of the instructions and data required to execute the task to the second processor core. The second processor core directly uses the storage addresses as prefetch information to obtain the instructions and data for executing the task from the memory. The second processor core can obtain the prefetch information for executing the task without performing a prefetch operation, reducing the area, cost, and power consumption of the devices used to perform the prefetch operation in the second processor core. At the same time, the area, cost, and power consumption of the second processor core are reduced. When the second processor core is integrated into the chip, the area and cost of the chip can be reduced. On the other hand, based on the execution results of the two processor cores, the failure rate of at least one of the two processor cores can be verified, improving the security performance of the CPU. In addition, the address information is the storage address of the instructions and data required to execute the task. Using the address information as prefetch information (or speculative information) improves the accuracy rate of the prefetch information (speculative information). Further, the second processor core uses the instructions and data to execute the task, improving the efficiency of executing the task, enhancing the processing performance of the second processor core, and thus improving the overall performance of the second processor core.
[0007] In a possible implementation manner of the first aspect, the first processor core includes a cache queue; the first processor core sending address information to the second processor core includes: the cache queue sending address information to the second processor core. In the above possible implementation manner, using the cache queue to transmit address information improves the transmission efficiency.
[0008] In a possible implementation manner of the first aspect, the first processor core further includes a memory access control module; the method further includes: the memory access control module sending multiple storage addresses to the cache queue, the multiple storage addresses being the storage addresses accessed by the first processor core through the memory access control module during the execution of the task; the cache queue filtering out duplicate addresses among the multiple storage addresses to obtain address information. In the above possible implementation manner, the cache queue filtering out duplicate addresses among the multiple storage addresses improves the transmission efficiency of the address information. Further, when transmitting the address information to the second processor core, the storage capacity occupied by the cache address information in the second processor core is reduced.
[0009] In a possible implementation manner of the first aspect, the first processor core sending address information to the second processor core includes: the first processor core sending address information to the second processor core through a bus, where the bus is separately designed for transmitting address information. In the above possible implementation manner, using the separately designed bus to transmit address information improves the rate of transmitting address information.
[0010] In a possible implementation of the first aspect, the second processor core includes a cache; before the second processor core fetches instructions and data from the memory using the address information as prefetch information, the method further includes: the second processor core writes the address information into the cache, and writes the instructions and data corresponding to the address information into the cache. In the above possible implementation, the processor core can fetch instructions and data from the cache faster than from the memory, which improves the rate of fetching instructions and data and enhances the performance of the second processor core.
[0011] In a possible implementation of the first aspect, the method further includes: the verification module verifies the failure rate of at least one of the first processor core or the second processor core according to the execution result. In the above possible implementation, the problem of the processor failure rate is solved, and the security performance of the processor is improved.
[0012] In a second aspect, a processing device for processor failure rate is provided. The processing device includes a first processor core and a second processor core; the first processor core is configured to send address information to the second processor core, where the address information includes the storage addresses of the instructions and data used by the first processor core to execute tasks in the memory, and the performance of the first processor core is higher than that of the second processor core; the second processor core is configured to fetch instructions and data from the memory using the address information as prefetch information, and execute tasks according to the instructions and data to obtain an execution result, where the execution result is used to verify the failure rate of at least one of the first processor core or the second processor core.
[0013] In a possible implementation of the second aspect, the first processor core includes a cache queue, and the cache queue is configured to send address information to the second processor core.
[0014] In a possible implementation of the second aspect, the first processor core further includes a memory access control module; the memory access control module is configured to send multiple storage addresses to the cache queue, where the multiple storage addresses are the storage addresses accessed by the first processor core through the memory access control module during the execution of tasks; the cache queue is configured to filter out duplicate addresses from the multiple storage addresses to obtain address information.
[0015] In a possible implementation of the second aspect, the processing device further includes a bus; the first processor core is configured to send address information to the second processor core through the bus.
[0016] In a possible implementation of the second aspect, the second processor core includes a cache; the second processor core is further configured to: write the address information into the cache before fetching instructions and data from the memory using the address information as prefetch information.
[0017] In a possible implementation of the second aspect, the processing device further includes: a verification module, configured to verify the failure rate of at least one of the first processor core or the second processor core according to the execution result.
[0018] In a third aspect, an electronic device is provided. The electronic device includes a memory and a processing device. The memory is used to store computer instructions. The processing device includes a first processor core and a second processor core. The processing device is configured to execute the processing method of the processor failure rate provided in the first aspect or any possible implementation of the first aspect as described above.
[0019] In a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores computer instructions. When the computer instructions run on the processing device, the processing device is caused to execute the processing method of the processor failure rate provided in the first aspect or any possible implementation of the first aspect as described above.
[0020] In a fifth aspect, a computer program product including instructions is provided. When the computer program product runs on the processing device, the processing device is caused to execute the processing method of the processor failure rate provided in the first aspect or any possible implementation of the first aspect as described above.
[0021] It can be understood that the processing device, electronic device, computer-readable storage medium, and computer program product provided above can be used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods provided above, which will not be elaborated here. Description of the Drawings
[0022] Figure 1 It is a schematic structural diagram of a chip system;
[0023] Figure 2 It is a schematic structural diagram of another chip system;
[0024] Figure 3 It is a schematic structural diagram of yet another chip system;
[0025] Figure 4 It is a schematic diagram of the execution efficiency of a processor core provided by an embodiment of the present application;
[0026] Figure 5 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application;
[0027] Figure 6 It is a schematic structural diagram of a processor provided by an embodiment of the present application;
[0028] Figure 7 It is a flowchart of a processing method of a processor failure rate provided by an embodiment of the present application;
[0029] Figure 8 It is a flowchart of another method for processing the processor failure rate provided by the embodiment of the present application;
[0030] Figure 9 It is a schematic structural diagram of a device for processing the processor failure rate provided by the embodiment of the present application. Detailed implementation manners
[0031] In the present application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B may be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (item)" or its similar expression refers to any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a, b, or c may represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c may be single or multiple. In addition, the embodiments of the present application use words such as "first" and "second" to distinguish the same items or similar items with basically the same functions and roles. For example, the first threshold and the second threshold are only used to distinguish different thresholds and do not limit their sequence. Those skilled in the art can understand that the words such as "first" and "second" do not limit the quantity and execution order.
[0032] It should be noted that in the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, using words such as "exemplary" or "for example" aims to present relevant concepts in a specific manner.
[0033] Before introducing the embodiments of the present application, the relevant knowledge of dual-core lockstep is first introduced and explained.
[0034] Artificial intelligence technology has been widely applied in fields such as intelligent vehicles, robots, and aerospace. With the rapid development of artificial intelligence technology, users' requirements for the security level and performance of processors are getting higher and higher. A processor includes a central processing unit (CPU). As the operation and control core of a computer system, the CPU is the final execution unit for information processing and program running. However, during use, due to quality defects, the environment, aging, or improper use, etc., the failure rate of the CPU will occur, that is, the CPU will malfunction. Therefore, there is an urgent need for a method to solve the problem of the CPU failure rate and improve the performance of the CPU.
[0035] Currently, the problem of the CPU failure rate can be solved through dual core lock step (DCLS) to improve the performance of the CPU. DCLS is specifically implemented through the following three solutions.
[0036] Solution 1, as Figure 1 shown in Figure 1It is a schematic structural diagram of a processor 10, which includes CPU cores 110, 120, an input / output port 130, and a verification module 140, as well as a cache 150 and a main memory 160 corresponding to the CPU core 110, and a cache 170 and a main memory 180 corresponding to the CPU core 120. The input / output port 130 provides an interface between the processor 10 and a peripheral interface module. For example, the peripheral interface module can be a keyboard, a mouse, or a universal serial bus (USB) device, etc. The processor 10 can be used to write the instructions and data received by the input / output port 130 into the main memory. For example, the instructions and data can be pre-written by the user through the keyboard, and the main memory can be a DDR (double data date) memory. Before executing a task, each CPU core reads the instructions and data corresponding to the task from the corresponding main memory and writes the read instructions and data into the corresponding cache. The cache is, for example, an on-chip cache located on the same chip as the CPU. Each CPU core executes the task based on the instructions and data in the corresponding cache and writes the execution result into the corresponding main memory through a write request. The write request includes the execution result and an address, which is the address corresponding to the storage space in the main memory for storing the execution result. At the same time, the CPU core sends the write request to the verification module 140. The verification module 140 (such as a parallel comparator) compares whether the two most recently written execution results in the memory are consistent every clock cycle (such as 60 s). For example, the verification module 140 uses an error correction code (ECC) or error detection and correction (EDC) to compare whether the execution results are consistent. If the execution result of the CPU core 110 is inconsistent with the execution result of the CPU core 120, it indicates that a certain CPU core has a fault, and the fault is processed using a fault handling module, thereby solving the problem of CPU failure rate. Figure 1 In the shown example, dual-core lockstep is implemented by the CPU cores 110, 120, the cache 150, the cache 170, the main memory 160, and the main memory 180. Figure 1 Only a partial structure of the processor is shown.
[0037] Solution 2: As Figure 2 shown, Figure 2 It is a schematic structural diagram of another processor 20, which includes CPU cores 210, 220, an input / output port 230, a verification module 240, a main memory 250, a cache 260 corresponding to the CPU core 210, and a cache 270 corresponding to the CPU core 220. Among them, the main memory 250 is shared by the CPU cores 210 and 220. Figure 2In the example shown, dual-core lockstep is implemented by CPU core 210, CPU core 220, cache 260, and cache 270. The specific process of implementing dual-core lockstep is as follows: Figure 1 The process of implementing dual-core lockstep by the processor 10 shown in FIG. 1 is similar and will not be described in detail here.
[0038] Option 3: Figure 3 As shown, Figure 3 3 is a schematic diagram of the structure of another processor 30, which includes a CPU core 310, a CPU core 320, an input and output port 330, a check module 340, a main memory 350 and a cache 360. The main memory 350 and the cache 360 are shared by the CPU core 210 and the CPU core 220. Figure 3 In the example shown, dual-core lockstep is implemented by CPU core 310 and CPU core 320. The specific process of implementing dual-core lockstep is as follows: Figure 1 The process of implementing dual-core lockstep by the processor 10 shown in FIG. 1 is similar and will not be described in detail here.
[0039] Optional, above Figure 1 , Figure 2 and Figure 3 Some components of the processor shown (e.g., CPU core and cache) can be integrated on the chip. Figure 1 , Figure 2 and Figure 3 The processor shown may include more CPU cores, such as, Figure 1 The processor shown in may include 3 CPU cores, Figure 1 , Figure 2 and Figure 3 This is merely an example and is not intended to be limiting.
[0040] However, in the above scheme, in order to achieve dual-core lock-step, it is necessary to set up at least two identical CPU cores, and at least two identical CPU cores perform the same task. Whether the CPU core is faulty is determined based on the results of different CPU cores performing the same task (for example, message transmission, message modification, image denoising tasks or picture processing tasks, etc.). However, setting up dual-core lock-step (that is, setting up two identical CPU cores in one chip) increases the area and cost of the chip system.
[0041] Based on this, an embodiment of the present application provides a method for processing the processor failure rate. When multiple processor cores process the same service (or homogeneous service), this method can obtain the address information obtained by the high-performance processor core sharing the task execution among the cores, and execute the services of other cores based on the address information (which can be used as prefetch information or speculative information), improving the accuracy of the prefetch information, thereby improving the processing performance of other cores; verifying the implementation efficiency of at least one processor core among multiple processor cores based on the execution result to improve the security performance of the processor, thereby enhancing the overall performance of the processor.
[0042] Exemplarily, as Figure 4 shown, Figure 4 the arrows in represent processor cores, and the length of each arrow represents the execution efficiency of the processor core. Among them, in the first method, multiple processor cores with the same performance (for example, Figure 1 , Figure 2 and Figure 3 the CPU cores shown in) execute the same task, and the execution efficiencies of the multiple processor cores are the same; in the second method, a first processor core with higher performance and other processor cores with performance lower than the first processor core execute the same task. In the first stage, the execution efficiency of the first processor core is faster than that of other processor cores. At this time, the first processor core shares the address information with other cores; in the second stage, after other cores execute the service according to the shared address information, the finally achieved execution efficiency is close to that of the first processor core, thereby improving the processing performance of the processor.
[0043] The technical solution provided by the embodiment of the present application can be applied to an electronic device including a high-security-level processor. The electronic device can be, but is not limited to, a mobile robot, a drone, a vehicle-mounted device, an aerospace device, etc. The following will describe the structure of the electronic device in combination with Figure 5 Exemplarily, Figure 5 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. The electronic device 50 may include: a processor 510, a bus 520, a memory 530, and a communication interface 540. The processor 510, the memory 530, and the communication interface 540 are connected through the bus 520.
[0044] It should be understood that in this embodiment, the processor 510 is the control center of the electronic device 50, connecting various parts of the entire device through various interfaces and buses 520. By running or executing software programs and / or software modules stored in the memory 530, and by invoking the data stored in the memory 530, it executes various functions of the electronic device 50 and processes data, thereby exercising overall control over the electronic device 50. The processor 510 may be a CPU, and this processor 510 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The processor 510 may also be a graphics processing unit (GPU), a neural network processing unit (NPU), a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the program of the present application solution.
[0045] In the embodiment of the present application, the processor 510 may be a multi-core processor. Exemplarily, Figure 6 FIG. is a schematic structural diagram of a processor 510 provided for the embodiment of the present application, such as Figure 6As shown, the processor 510 may include multiple processor cores (also referred to as cores). For example, the processor 510 may include a first processor core and a second processor core. The first processor core includes a memory access control module and a cache queue. The memory access control module may also be referred to as a load store unit (LSU), and the cache queue may be a prefetch input buffer (inbuf). Further, the first processor core may also include a level 1 cache (L1), a level 2 cache (L2), a hardware prefetch (HWP), a prefetch to level 1 cache (PFL1) interface, a prefetch to level 2 cache (PFL2) interface, a prefetch to level 3 cache (PFL3) interface, and other modules. In the figure, the first processor core including LSU, inbuf, L1, L2, HWP, PFL1 interface, PFL2 interface, and PFL3 interface is taken as an example for illustration, and LSU and L1 are represented as LSU+L1. Optionally, as Figure 6 shown, other processor cores (slave cores), such as the second processor core, may have a structure similar to that of the first processor core. Figure 6 In the example where the processor 510 includes a first processor core and a second processor core, it does not limit the number of processor cores.
[0046] Optionally, the processor 510 may include a bus and a verification module. The bus can be used to transmit address information. Among them, the bus can be a multiplexed bus or a newly added bus. The verification module can be used to compare whether the execution results (output results) of the first processor core and the second processor core are consistent. For example, the verification module can be a parallel comparator. Figure 6 Only a partial structure of the processor 510 is schematically shown in the figure. Those skilled in the art can understand that Figure 6 the structure of the processor shown in the figure does not limit the processor, and it may include more or fewer components than shown, or combine some components, or have different component arrangements.
[0047] In a possible embodiment, Figure 6 the processor 510 shown in the figure may be applied in an electronic device in the form of a chip. For example, the processor cores and storage included in the processor 510 may be provided on the chip.
[0048] Continue to refer to Figure 5, the communication interface 540 is used to implement the communication between the electronic device 50 and external devices or components.
[0049] The bus 520 may include a path for transmitting information between the above components (such as the processor 510 and the memory 530). The bus 520 may include an address bus, a data bus, a control bus, etc. However, for the sake of clarity, all kinds of buses are labeled as the bus 520 in the figure. The bus 520 may be a Peripheral Component Interconnect Express (PCIe) bus, or an Extended Industry Standard Architecture (EISA) bus, a Unified Bus (Ubus or UB), a Compute Express Link (CXL), a Cache Coherent Interconnect for Accelerators (CCIX), etc.
[0050] It is worth noting that Figure 5 only takes the example that the electronic device 50 includes 1 processor (for example, the processor 510) and 1 memory (for example, the memory 530). Here, the processor 510 and the memory 530 are respectively used to indicate a type of device or component. In specific embodiments, the number of each type of device or component can be determined according to service requirements.
[0051] The memory 530 can be used to store data, software programs, and software modules; it mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function, such as a sound playback function or an image playback function, etc.; the data storage area can store data created according to the use of the electronic device, such as audio data, image data, or table data, etc. For example, the memory 530 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchlink dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM).
[0052] Although not shown, the electronic device may further include an audio component and a communication component, etc. For example, the audio component includes a microphone, and the communication component includes a wireless fidelity (WiFi) module or a Bluetooth module, etc. Details are not described herein in the embodiments of the present application. Those skilled in the art can understand that Figure 5 the structure of the electronic device shown does not limit the electronic device, and it may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0053] Before introducing the method for processing the processor failure rate provided by the embodiments of the present application, first, the application scenario of the method for processing the processor failure rate provided by the embodiments of the present application is introduced and described. The method for processing the processor failure rate provided by the embodiments of the present application is applied to the aboveFigure 5 In the processor 510 shown. In scenarios with high security level requirements, the first processor core and the second processor core in the processor 510 can be used to execute the same task simultaneously. Since the performance of the first processor core is higher / stronger than that of the second processor core, that is, the execution efficiency of the first processor core for executing tasks is higher than that of the second processor core for executing tasks, the first processor core can send the storage addresses of the instructions and data used to execute the task in the memory to the second processor core. The second processor core can use the storage addresses as prefetch information (or speculative information) of the task to execute the task, and send a request carrying the execution result to the verification module. This request is used to indicate writing the execution result into the memory 530. The verification module verifies whether the execution results of the first processor core and the second processor core are consistent to determine the failure rate of at least one of the first processor core and the second processor core, and uses the fault handling module to perform fault handling to solve the problem of the processor failure rate and improve the security performance of the processor.
[0054] Next, in combination with Figure 6 the processor shown, the method for handling the processor failure rate provided by the embodiments of the present application will be introduced and described. The processor may include multiple processor cores, and the multi-core processor may include a first processor core and a second processor core. The processor may include a CPU, a GPU, an NPU, etc., and the embodiments of the present application do not make specific limitations in this regard.
[0055] Figure 7 is a schematic flowchart of a method for handling the processor failure rate provided by the embodiments of the present application. The method includes:
[0056] S701: The first processor core sends address information to the second processor core. The address information includes the storage addresses of the instructions and data used by the first processor core to execute the task in the memory, and the performance of the first processor core is higher than that of the second processor core.
[0057] Among them, the first processor core is the processor core with higher performance among multiple processor cores, that is, the first processor core is the processor core with higher execution efficiency among multiple processor cores. The first processor core can also be called the main core or the leading core. The second processor core is any one of the processor cores with performance inferior to that of the first processor core among multiple processor cores, that is, the second processor core is any one of the processor cores with execution efficiency inferior to that of the first processor core among multiple processor cores. The second processor core can also be called the slave core or the following core. The memory can be the memory 530 shown above Figure 5 In practical applications, the first processor core can be a large core with more hardware resources, and the second processor core can be a small core with fewer hardware resources, that is, a simplified small core.
[0058] Optionally, the first processor core as the leading core can be pre-set or specified. Exemplarily, the leading core is determined according to the physical locations of multiple processor cores. For example, the first core in terms of physical location is set as the leading core; alternatively, the leading core is specified by software running on the processor. The embodiments of the present application do not specifically limit the manner of specifying the leading core.
[0059] Secondly, the task is the same task executed by the first processor core and the second processor core. The same task means that the data and / or instructions required to be accessed during the running of the two tasks are the same. For example, the running instructions, data, or branches are all the same. In practical applications, the first processor core and the second processor core are used to execute the same task simultaneously, that is, the first processor core and the second processor core adopt dual-core lockstep to solve the problem of processor inefficiency. The first processor core can send the address information to the second processor core after determining the address information actually used for task execution. Specifically, the address information can be sent during or after the execution of the task. This embodiment does not limit this.
[0060] In a possible embodiment, when the first processor core receives an instruction to execute a task, the first processor core executes the task and accesses a memory or a cache (such as a level-1 cache or a level-2 cache) through a memory access control module to obtain the instructions and data required for task execution, and obtains the address information; similarly, when the second processor core receives an instruction to execute a service, the second processor core executes the service. Among them, the running speed of the first processor core is greater than that of the second processor core, that is, the running speed of the first processor core is faster than the running progress of the second processor core. In this way, the first processor core sending the address information to the second processor core can improve the speculative performance of the second processor core, and further improve the performance of the second processor core in processing tasks.
[0061] In addition, the address information can include instruction address information and data address information. The instruction address information is the storage address of the instructions required for the first processor core to execute the task in the memory. The instructions required for the first processor core to execute the task can include, but are not limited to, front-end instruction fetching instructions, jump instructions (including jump directions and jump destinations), back-end memory access instructions, sequential instructions, and value prediction instructions, etc.; the data address information is the storage address of the data required for the first processor core to execute the task.
[0062] In a possible embodiment, before the first processor core executes a task, the hardware prefetcher in the first processor core can be used to write the instructions and data required for executing the task into the cache corresponding to the first processor core (such as L1 and L2) in units of cache lines in advance. A cache line includes instructions and data, as well as the storage address corresponding to the instructions and data in the memory. A cache line is usually 512 bits. During the process of the first processor core executing the task, the memory access control module sends an access request to the memory. The access request includes the storage address in the memory, and the access request is usually 8 bits. The memory access control module (such as LUS) accesses the memory indirectly through the cache. Each time it accesses the memory, it needs to traverse the cache lines in the cache to find whether there is a storage address to be accessed in the cache line. If there is a storage address in the cache line and both the instructions and data corresponding to the storage address are valid, the instructions and data stored in the cache line are directly read; if there is no storage address in the cache line, or there is a storage address in the cache line but the instructions or data corresponding to the storage address are invalid, the instructions and data are directly read from the memory into the cache, and then the instructions and data are read from the cache. During this process, a cache line can be accessed by the memory access control module multiple times, that is, a storage address can be accessed multiple times.
[0063] In a possible instance, after the first processor core determines multiple storage addresses, the memory access control module (such as LUS) sends the multiple storage addresses to a cache queue (such as inbuf). The multiple storage addresses are the storage addresses accessed through the memory access control module during the process of the first processor core executing the task. The cache queue stores and filters the duplicate addresses among the multiple storage addresses to obtain address information, and the cache queue sends the address information to the second processor core. For example, the cache queue sends the address information through the PFL3 interface. Among them, the cache queue can be a multi-input and single-output cache queue, which can be used not only to filter duplicate addresses but also to balance the bandwidth difference between input and output to improve the transmission efficiency of the address information. Therefore, as Figure 8 shown, the method provided by the embodiment of the present application further includes:
[0064] S01: The memory access control module sends multiple storage addresses to the cache queue. The multiple storage addresses are the storage addresses accessed through the memory access control module during the process of the first processor core executing the task; the cache queue filters the duplicate addresses among the multiple storage addresses to obtain address information.
[0065] Step S701 thus includes: The cache queue sends the address information to the second processor core.
[0066] Exemplarily, the address information can be the memory access request address with a first-level cache miss, or a subset or the whole set of any regular first-level high-speed memory access addresses, etc. The embodiment of the present application does not make specific limitations on this.
[0067] In a possible embodiment, the second processor core receives address information from the first processor core. Optionally, the multiple processor cores are coupled through a bus, which can be the original bus in the processor (i.e., the bus in the processor is reused), or a bus specifically set in this application (for example, a newly added bus), and the embodiments of this application do not make specific limitations on this. Therefore, in the method provided by the embodiments of this application, step S701 specifically includes: the first processor core sends address information to the second processor core through the bus. In this embodiment, by transmitting address information through a specifically set bus, the transmission efficiency of address information is improved.
[0068] Optionally, as Figure 6 shown, the structure of the second processor core is similar to that of the first processor core. The second processor core may include: a memory access control module and a cache queue, and the cache queue may be a prefetch input buffer inbuf; further, the second processor core may also include other modules such as a level 1 cache L1, a level 2 cache L2, a hardware prefetch unit HWP, a PFL1 interface, a PFL2 interface, and a PFL3 interface. In addition, the PFL3 interfaces of the second processor core and the first processor core are coupled through a bus.
[0069] In a possible example, in combination with Figure 6 shown, the second processor core receives address information from the first processor core through the bus. Specifically, the hardware prefetch unit (for example, HWP) in the second processor core receives address information from the first processor core through the bus.
[0070] It can be understood that the multiple processor cores may include multiple slave cores, that is, in addition to the first processor core, the multiple processor cores may also include other slave cores similar to the first processor core, and the first processor core (i.e., the leading core) may send address information to each slave core. Figure 6 In
[0071] Furthermore, the method provided by the embodiments of this application further includes:
[0072] S02: The second processor core writes the address information into the cache corresponding to the second processor core.
[0073] S702: The second processor core obtains instructions and the data from the memory by using the address information as prefetch information.
[0074] In a possible embodiment, the second processor core may include a cache and a hardware prefetcher (e.g., HWP). The hardware prefetcher can be used to receive address information, convert the received address information into the corresponding physical address of the memory, and use the physical address as prefetch information to obtain the instructions and data corresponding to the task from the memory. The hardware prefetcher is also used to write the physical address, as well as the obtained instructions and data, into the cache (e.g., L1) corresponding to the second processor core.
[0075] Optionally, since the address information may include instruction address information and data address information, the corresponding physical address obtained by converting the address information may also include an instruction physical address and a data physical address. Among them, the instruction physical address is the storage address of the instruction corresponding to the task, and the data physical address is the storage address of the data corresponding to the task. The instruction physical address and the data physical address can be stored in different caches respectively. For example, the cache corresponding to the second processor core may include a first cache and a second cache. The first cache can be used to store instructions, and the second cache can be used to store data. The instruction physical address can be stored in the first cache. Optionally, the first cache is also used to store the instructions corresponding to the task; the data physical address can be stored in the second cache. Optionally, the second cache can be used to store the data corresponding to the task.
[0076] Among them, the cache can be any one of the first-level cache (L1), second-level cache (L2), or third-level cache (L3) corresponding to the second processor core. All the data or instructions stored in each level of cache are part of the next-level cache. The closer the cache is to the second processor core, the faster and smaller it is. For example, the first-level cache is adjacent to the second processor core, and the L1 cache is the cache with the smallest capacity and the fastest read / write speed among the three-level caches. The second processor core can write the address information into the first-level cache, second-level cache, or third-level cache. The embodiments of the present application do not make specific limitations on this. Figure 6 Taking the processor shown in which the second processor core includes L1 and L2 as an example.
[0077] Furthermore, the method provided by the embodiments of the present application further includes:
[0078] S703: The second processor core executes the task according to the instructions and data to obtain an execution result, and the execution result is used to verify the failure rate of at least one of the first processor core or the second processor core.
[0079] Specifically, during the process of the second processor core executing the task, it directly obtains the instructions and data for executing the task from the cache corresponding to the second processor core as prefetch information, executes the task according to the obtained instructions and data to obtain an execution result, and the second processor core sends the execution result to a verification module (e.g., a parallel comparator).
[0080] Further, the method provided by the embodiment of the present application further includes:
[0081] S704: The verification module verifies the failure rate of at least one of the first processor core or the second processor core according to the execution result.
[0082] Further, the second processor core writes the execution result into the memory through a first write request. The first write request includes the execution result of the second processor core and a first address, and the first address is the address corresponding to the storage space in the memory for storing the execution result of the second processor core. The second processor core simultaneously sends the first write request to the verification module. Similarly, the first processor core writes the execution result into the memory through a second write request. The second write request includes the execution result of the first processor core and a second address, and the second address is the address corresponding to the storage space in the memory for storing the execution result of the first processor core. The first processor core simultaneously sends the second write request to the verification module. Among them, the first address and the second address may be the same address, that is, the first address and the second address can be used to indicate the same storage space in the memory.
[0083] The verification module compares whether the two most recently written execution results in the memory are consistent in each clock cycle (for example, 60s). For example, the verification module can use an error correction code (ECC) or error detection and correction (EDC) to verify whether the execution results (including data, address, and control lines) are consistent. If the execution results are consistent, the first processor core and the second processor core are normal; if the execution results are inconsistent, at least one of the first processor core and the second processor core is abnormal (including the first processor core is abnormal, or the second processor core is abnormal, or both the first processor core and the second processor core are abnormal), a fault signal is sent, and the fault processing module is used for fault processing, thereby solving the problem of the processor failure rate and improving the security performance of the processor.
[0084] In the embodiment of the present application, when the first processor core and the second processor core process the same service, the first processor core sends address information to the second processor core, and the second processor core directly uses the address information as prefetch information (or called speculative information) to obtain the instructions and data for executing tasks from the memory. The second processor core can obtain the prefetch information without performing a prefetch operation, reducing the area overhead of the hardware prefetching unit. On the other hand, the accuracy of the prefetch information is improved, thereby improving the processing performance of the second processor; in addition, the failure rate of at least one of the two processor cores is verified by using the execution result, that is, CPU faults are detected. The security performance of the processor core is improved.
[0085] Further, the above embodiments are described by taking the operating speed of the first processor core being higher than that of the second processor core as an example. In fact, this setting is not used to limit the application scenarios of the embodiments. In practical applications, multiple processor cores may have the same or different capabilities or processing speeds, and it is not limited which one has stronger capabilities. As long as the operating information obtained after one of the processor cores executes a task can be shared with another processor core to execute the technical solutions mentioned in this embodiment. For example, still taking Figure 6 as an example, the functions of the leading core and the slave core can be swapped. The second processor core (slave core) can send its operating information to the first processor core (leading core) so that the first processor core can execute a process similar to Figure 7 the embodiment process described above to achieve a similar function. For example, the second processor core may execute at least part of the functions of a task first due to software scheduling or user selection and obtain address information, and this address information can be shared with the first processor core to continue executing the same task in order to achieve a similar effect. It can be understood that the leading core and the slave core mentioned in this embodiment, as well as the differences between these two types of cores, are only an applicable scenario, but are not used to limit the technical solutions.
[0086] Based on this, the embodiment of the present application also provides a processing device for the processor failure rate, as Figure 9 described, and this device can be applied to a processor. As Figure 9 shown, the device includes a first processor core and a second processor core. In the embodiment of the present application, the first processor core can be used to execute the steps of S701 and S01 in the above method embodiment, and / or other steps described herein, etc.; the second processor core can be used to execute the steps of S02, S702, and S703 in the above method embodiment, and / or other steps described herein, etc.
[0087] The verification module verifies the failure rate of at least one of the first processor core or the second processor core according to the execution result.
[0088] It can be understood that regarding the specific structures of the first processor core and the second processor core, as well as all relevant contents of each step involved in the above method embodiment, they can all be cited into the embodiment of this service processing device, and the embodiment of the present application will not elaborate herein.
[0089] On the other hand of the present application, an electronic device is also provided. This electronic device includes a memory and at least one processor. The memory is used to store computer instructions, and the at least one processor includes multiple processor cores and is used to execute the computer instructions so that the electronic device can implement the processing method for the processor failure rate provided above. Optionally, the at least one processor includes the processing device for the processor failure rate provided above.
[0090] It is understandable that all relevant content of each step involved in the above method embodiments can be cited in the embodiments of the processing device for the processor failure rate and the embodiments of this electronic device. The embodiments of this application will not be elaborated herein.
[0091] In several embodiments provided by this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed.
[0092] The units described as separate components may or may not be physically separated. The components shown as units may be one physical unit or multiple physical units, that is, they can be located in one place or distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0093] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. The readable storage medium may include: various media such as USB flash drives, mobile hard disks, read-only memories, random access memories, magnetic disks, or optical discs that can store program codes. Based on such an understanding, the technical solution of the embodiments of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product.
[0094] In another embodiment of this application, a readable storage medium is further provided. Computer-executable instructions are stored in the readable storage medium. When a device (which can be a single-chip microcomputer, a chip, etc.) or a processor executes the steps in the above method embodiments.
[0095] In yet another embodiment of this application, a computer program product is further provided. The computer program product includes computer instructions, and the computer instructions are stored in a readable storage medium; at least one processor of the device can read the computer instructions from the readable storage medium, and at least one processor executes the computer instructions to enable the device to perform the steps in the above method embodiments.
[0096] Finally, it should be noted that the above are only the specific implementation manners of this application, but the protection scope of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A method for processing the processor failure rate, characterized in that The method includes: The first processor core sends address information to the second processor core, where the address information includes the storage addresses in the memory of the instructions and data used by the first processor core to execute a task, and the performance of the first processor core is higher than that of the second processor core; The second processor core obtains the instructions and the data from the memory by using the address information as prefetch information; The second processor core executes the task according to the instructions and the data, and obtains an execution result, where the execution result is used to verify the failure rate of at least one of the first processor core and the second processor core.
2. The method according to claim 1, wherein The first processor core includes a cache queue; The first processor core sending the address information to the second processor core includes: The cache queue sends the address information to the second processor core.
3. The method according to claim 2, wherein The first processor core further includes a memory access control module; The method further includes: The memory access control module sends a plurality of storage addresses to the cache queue, where the plurality of storage addresses are the storage addresses accessed by the first processor core through the memory access control module during the execution of the task; The cache queue filters out duplicate addresses among the plurality of storage addresses to obtain the address information.
4. The method according to any one of claims 1 to 3, characterized in that The first processor core sending the address information to the second processor core includes: The first processor core sends the address information to the second processor core through a bus.
5. The method according to any one of claims 1-4, characterized in that, The second processor core includes a cache; Before the second processor core obtains the instructions and the data from the memory by using the address information as prefetch information, the method further includes: The second processor core writes the address information into the cache.
6. The method according to any one of claims 1-5, characterized in that, The method further includes: A verification module verifies the failure rate of at least one of the first processor core and the second processor core according to the execution result.
7. A processing device for processor failure rate, characterized in that The processing device includes a first processor core and a second processor core; The first processor core is configured to send address information to the second processor core, where the address information includes the storage addresses in the memory of the instructions and data used by the first processor core to execute a task, and the performance of the first processor core is higher than that of the second processor core; The second processor core is configured to obtain the instructions and the data from the memory by using the address information as prefetch information, and execute the task according to the instructions and the data, and obtain an execution result, where the execution result is used to verify the failure rate of at least one of the first processor core and the second processor core.
8. The processing device according to claim 7, wherein The first processor core includes a cache queue, The cache queue is configured to send the address information to the second processor core.
9. The processing device according to claim 8, characterized in that, The first processor core further includes a memory access control module; The memory access control module is configured to send a plurality of storage addresses to the cache queue, where the plurality of storage addresses are the storage addresses accessed by the first processor core through the memory access control module during the execution of the task; The cache queue is configured to filter out duplicate addresses among the plurality of storage addresses to obtain the address information.
10. The processing device according to any one of claims 7-9, characterized in that, The processing device further includes a bus; The first processor core is configured to send the address information to the second processor core via the bus.
11. The processing device according to any one of claims 7-10, characterized in that, The second processor core includes a cache; The second processor core is further configured to: write the address information into the cache before obtaining the instruction and the data from the memory by using the address information as prefetch information.
12. The processing device according to any one of claims 7-11, characterized in that, The processing device further includes: a verification module, configured to verify the failure rate of at least one of the first processor core and the second processor core according to the execution result.
13. An electronic device, characterized in that, The electronic device includes a memory and a processing device, the memory is configured to store computer instructions, the processing device includes a first processor core and a second processor core, and the processing device is configured to execute the computer instructions to implement the method for processing the processor failure rate according to any one of claims 1-6.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and when the computer instructions run on the processing device, the processing device is caused to execute the method according to any one of claims 1-6.
15. A computer program product comprising instructions, characterized in that, When the computer program product runs on the processing device, the processing device is caused to execute the method according to any one of claims 1-6.
Citation Information
Cited By
Processor failure rate processing method and processing apparatus, and electronic device
WO2025130282A1