Processing tasks in a processing system

CN114625577BActive Publication Date: 2026-09-29IMAGINATION TECH LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111491368.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-06-29
Filing Date
2021-12-08
Publication Date
2026-09-29
Estimated Expiration
2041-12-08

AI Technical Summary

Technical Problem

然而,实现锁步处理器的面积和功率消耗(以及因此成本)的增加可能是不可接受的或在这些应用中是不期望的

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114625577B_ABST
    Figure CN114625577B_ABST
Patent Text Reader

Abstract

A method of processing an input task in a processing system, the method comprising: replicating the input task to form a first task and a second task; allocating memory, the memory comprising: a first memory block configured to store read-write data to be accessed during processing of the first task; a second memory block configured to store a copy of the read-write data to be accessed during processing of the second task; and a third memory block configured to store read-only data to be accessed during processing of both the first task and the second task; and processing the first task and the second task at processing logic of the processing system to generate a first output and a second output, respectively.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to processing tasks in a processing system. Background Technology

[0002] This disclosure relates to a processing system and a method for processing tasks in the processing system.

[0003] In safety-critical systems, at least some components must meet safety objectives sufficient to enable the system as a whole to meet the level of safety considered necessary for the system. For example, in most jurisdictions, seatbelt retractors in vehicles must meet specific safety standards so that vehicles equipped with such devices can pass safety tests. Similarly, vehicle tires must meet specific standards so that vehicles equipped with such tires can pass safety tests appropriate for a particular jurisdiction. Safety-critical systems are typically those whose failure would significantly increase the risk to human safety or the environment.

[0004] Processing systems, such as data processing devices, often constitute components of safety-critical systems, either as dedicated hardware or as processors running safety-critical software. For example, fly-by-wire flight systems in aircraft, pilot assistance systems, railway signaling systems, and control systems for medical devices are typically safety-critical systems operating on data processing devices. When a data processing device is part of a safety-critical system, it is necessary for the data processing device itself to meet safety objectives so that the system as a whole can meet an appropriate safety level. In the automotive industry, the safety level is typically the Automotive Safety Integrity Level (ASIL) as defined in the functional safety standard ISO 26262.

[0005] Increasingly, data processing devices used in safety-critical systems include processors running software. Both hardware and software components must meet specific safety objectives. Some software failures may be systematic failures caused by programming errors or poor error handling. These problems can usually be addressed through rigorous development practices, code reviews, and testing protocols. Even if a safety-critical system may not contain any systematic errors, random errors can still be introduced into the hardware, for example, due to transient events such as ionizing radiation, voltage spikes, or electromagnetic pulses. In binary systems, transient events can cause random bit flips in memory and along the processor's data path. Hardware may also have permanent failures.

[0006] The security objectives of a data processing device can be expressed as a set of metrics, such as the maximum number of failures within a given time period (typically expressed as failure-in-time or FIT), and the effectiveness of mechanisms for detecting single points of failure (Single Point Failure Mechanism or SPFM) and for detecting potential failures (Least Point Failure Mechanism or LFM). Various methods exist for achieving the security objectives set for the data processing device: for example, by providing hardware redundancy so that if one component fails, another can be used to perform the same task, or by using check data (e.g., parity bits or error correction codes) to enable the hardware to detect and / or correct minor data corruption.

[0007] For example, such as Figure 1 As shown, the data processor can be provided in a dual-lockstep arrangement 100, where a pair of identical processing units 101 and 102 are configured to process instruction streams 103 in parallel. Processing units 101 and 102 are typically synchronized for each instruction stream, such that both processing units 101 and 102 execute the instruction stream simultaneously, loop by loop. The output of either processing unit 101 or 102 can be used as the output 104 of the lockstep processor. A safety-critical system may fail if the outputs of processing units 101 and 102 do not match. However, due to the need for a second processing unit, the dual-lockstep processor will inevitably consume twice the chip area and approximately twice the power compared to a conventional processor.

[0008] In another example, by adding additional processor units (not shown) to the lockstep processor 100, error-free output can continue to be provided even if a fault is detected in one of these processor units. This can be achieved using a process called modular redundancy. Here, the output of the lockstep processor can be the output provided by two or more of its processing units, where the output of processing units that do not match the other units are ignored. However, this further increases the processor's area and power consumption.

[0009] Advanced driver assistance systems and autonomous vehicles may have data processing systems that must meet specific safety objectives. For example, autonomous vehicles must process large amounts of data in real time (e.g., data from radar, lidar, map data, and vehicle information) to make safety-critical decisions. Such safety-critical systems in autonomous vehicles typically need to meet the most stringent ASIL Level D of ISO 26262. However, the increased area and power consumption (and therefore cost) of implementing a lockstep processor may be unacceptable or undesirable in these applications. Summary of the Invention

[0010] This summary is provided to introduce, in a simplified form, a series of concepts further described below in the detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter.

[0011] According to a first aspect, a method for processing an input task in a processing system is provided, the method comprising: copying the input task to form a first task and a second task; allocating memory, the memory comprising: a first memory block configured to store read-write data to be accessed during processing of the first task; a second memory block configured to store a copy of the read-write data to be accessed during processing of the second task; and a third memory block configured to store read-only data to be accessed during both processing of the first task and the second task; and processing the first task and the second task at processing logic of the processing system to generate a first output and a second output, respectively.

[0012] The method may further include: forming a first signature and a second signature, the first signature and the second signature being features of a first output and a second output, respectively; comparing the first signature and the second signature; and issuing a fault signal if the first signature and the second signature do not match.

[0013] The first signature and the second signature, which are the characteristics of the first output and the second output respectively, may include one or more of the following: checksum, cyclic redundancy check, hash and fingerprint, determined on the first processed output and the second processed output respectively.

[0014] The method may further include forming a first signature and a second signature before the memory level of the first and second output access processing system.

[0015] The method may further include: before processing the first task and the second task, storing the read / write data at a memory address in the first memory block, and storing a copy of the read / write data at the corresponding memory address in the second memory block.

[0016] A first memory block and a second memory block can be allocated in the memory heap, with each memory address of the second memory block offset from the corresponding memory address in the first memory block by a fixed memory address step.

[0017] Multiple input tasks can be processed at the processing system, and for each pair of first and second tasks formed by the corresponding input tasks, the fixed memory address stride can be the same.

[0018] A memory heap can be a contiguous block of memory reserved for storing data to process one or more input tasks at the processing system. The memory heap is located in the memory of the processing system.

[0019] The method may further include: receiving a second output; identifying a reference to a memory address in a first memory block in the second output; updating the reference using a memory address step; and accessing the corresponding memory address in a second memory block using the updated reference.

[0020] The method may further include receiving an output and identifying that the output is received from a second task, so as to identify the output as a second output.

[0021] A third memory block can be allocated in the memory heap.

[0022] The method may also include simultaneously submitting a first task and a second task to the processing logic.

[0023] The method may further include: extracting data from a first memory block, a second memory block, and a third memory block into a cache, the cache being configured to be accessed by processing logic during the processing of a first task and a second task.

[0024] The input task can be a security task that is processed according to a predefined security level.

[0025] The processing logic may include a first processing element and a second processing element, wherein processing the first task and the second task at the processing logic of the processing system includes processing the first task at the first processing element and processing the second task at the second processing element.

[0026] The input task may be a test task, which includes a predefined set of instructions for execution on the processing logic. The predefined set of instructions is configured to perform a predetermined set of operations on the processing logic when executed against predefined input data. The method may also include receiving the test task at a processing unit including a first processing element and a second processing element.

[0027] The processing logic may include specific processing elements, wherein processing the first task and the second task at the processing logic of the processing system includes processing the first task at the specific processing element and processing the second task at the specific processing element.

[0028] The first and second outputs may include intermediate outputs generated during the processing of the first and second tasks, respectively. Intermediate outputs may be one or more of load, store, or atomic instructions generated during task processing.

[0029] The processing logic can be configured to handle the first and second tasks independently.

[0030] The input task can be a computation workgroup that includes one or more computational work items.

[0031] The method may further include, during the processing of the first task: reading read / write data from the first memory block; modifying the data according to the first task; and writing the modified data back to the first memory block.

[0032] The method may further include, during the processing of a second task: reading read / write data from a second memory block; modifying the data according to the second task; and writing the modified data back to the second memory block.

[0033] According to a second aspect, a processing system is provided, the processing system being configured to process input tasks, the processing system comprising: a task copying unit configured to copy the input tasks to form a first task and a second task; a memory allocation unit configured to allocate memory, the memory comprising: a first memory block configured to store read / write data to be accessed during processing of the first task; a second memory block configured to store a copy of the read / write data to be accessed during processing of the second task; and a third memory block configured to store read-only data to be accessed during both processing of the first task and the second task; and processing logic configured to process the first task to generate a first output, and to process the second task to generate a second output.

[0034] The processing system may further include: an inspection unit configured to form a first signature and a second signature, the first signature and the second signature being features of a first output and a second output, respectively; and a fault detection unit configured to compare the first signature and the second signature, and to issue a fault signal if the first signature and the second signature do not match.

[0035] The processing system may further include a memory stack, which includes a first memory block, a second memory block, and a third memory block.

[0036] The processing system described herein may be contained in hardware on an integrated circuit. A method for manufacturing the processing system as described herein at an integrated circuit manufacturing system may be provided. An integrated circuit definition dataset may be provided, which, when processed in an integrated circuit manufacturing system, configures the system to manufacture the processing system described herein. A non-transitory computer-readable storage medium may be provided, on which a computer-readable description of the processing system described herein is stored, which, when processed in an integrated circuit manufacturing system, causes the integrated circuit manufacturing system to manufacture an integrated circuit containing the processing system described herein.

[0037] An integrated circuit manufacturing system may be provided, comprising: a non-transitory computer-readable storage medium storing a computer-readable description of the processing system described herein; a layout processing system configured to process the computer-readable description to generate a circuit layout description of an integrated circuit including the processing system described herein; and an integrated circuit generation system configured to manufacture the processing system described herein based on the circuit layout description.

[0038] Computer program code for performing any of the methods described herein may be provided. A non-transitory computer-readable storage medium may be provided, on which computer-readable instructions are stored, which, when executed at a computer system, cause the computer system to perform any of the methods described herein.

[0039] As will be apparent to those skilled in the art, the above features can be appropriately combined, and can be combined with any aspect of the examples described herein. Attached Figure Description

[0040] The example will now be described in detail with reference to the accompanying drawings, in which:

[0041] Figure 1 A typical double-lockstep processor is shown.

[0042] Figure 2 The graphics processing unit configured according to the principles described herein is shown.

[0043] Figure 3 A data processing system including a graphics processing unit configured according to the principles described herein is shown.

[0044] Figure 4 It shows Figure 3 The diagram illustrates an exemplary logical arrangement of units in a data processing system, which are used to process input tasks according to the principles described herein.

[0045] Figure 5 A method for processing input tasks at a data processing system based on the principles described herein is illustrated.

[0046] Figures 6a to 6c An exemplary allocation of the first, second, and third memory blocks according to the principles described herein is shown.

[0047] Figure 7 An exemplary set of steps performed by the inspection and filtering unit according to the principles described herein is shown.

[0048] Figure 8 An integrated circuit manufacturing system for producing integrated circuits containing a data processing system is shown.

[0049] The accompanying drawings illustrate various examples. Those skilled in the art will understand that the element boundaries (e.g., boxes, groups of boxes, or other shapes) shown in the drawings represent one example of a boundary. In some examples, it may be that one element can be designed as multiple elements, or multiple elements can be designed as one element. Where appropriate, common reference numerals are used throughout the drawings to indicate similar features. Detailed Implementation

[0050] The following description is given by way of example to enable those skilled in the art to make and use the invention. The invention is not limited to the embodiments described herein, and various modifications to the disclosed embodiments will be readily apparent to those skilled in the art.

[0051] The embodiments will now be described by way of example only.

[0052] This disclosure relates to processing tasks at a processing system. The processing system may be referred to herein as a data processing system. A data processing system configured according to the principles herein can have any suitable architecture; for example, a data processing system can be operable to perform any kind of graphics, image, or video processing, general processing, and / or any other type of data processing.

[0053] A data processing system includes processing logic, which comprises one or more processing elements. For example, a data processing system may include multiple processing elements, which can be, for example, any kind of graphics and / or vector and / or stream processing elements. Each processing element may be a different physical core of a graphics processing unit (GPU) that constitutes the data processing system. Nevertheless, it should be understood that the principles described herein can be applied to processing elements of any suitable type of processing unit (e.g., a central processing unit (CPU) with a multi-core arrangement). Data processing systems can be applied to general computing tasks, particularly those that can be easily parallelized. Examples of general computing applications include signal processing, audio processing, computer vision, physics simulation, statistical computation, neural networks, and encryption.

[0054] A task can be any part of the work being processed at a processing element. For example, a task can define one or more processing actions to be performed on any kind of data (e.g., vector data) that a processing element of a data processing system can be configured to process. A data processing system can be configured to operate on multiple different types of tasks. In some architectures, different processing elements or groups of processing elements can be assigned to handle different types of tasks.

[0055] In the example, the task to be processed at the data processing system can be a computation workgroup comprising one or more computational work items. A computational work item can be an instance of a computation kernel (e.g., a computation shader). One or more computational work items can collaboratively operate on common data. The one or more computational work items can be grouped together into a so-called computational workgroup. Each computational work item in a computational workgroup can execute the same computation kernel (e.g., a computation shader), but each work item can operate on different portions of the common data shared by these work items. Such computational workgroups comprising one or more computational work items can be assigned to processing elements of the data processing system. Each computational workgroup can be independent of any other workgroup. In another example, the task to be processed at the data processing system can be a test task, as will be described in further detail herein.

[0056] Figure 2 A graphics processing unit configured according to the principles described herein is illustrated. It should be understood that although this disclosure will be described with reference to a data processing system including a graphics processing unit (GPU), the principles described herein can be applied to data processing systems including any suitable type of processing unit (e.g., a central processing unit (CPU) with a multi-core arrangement).

[0057] A graphics processing unit (GPU) 200 may be part of a data processing system. The GPU 200 includes multiple processing elements 204, labeled PE0 to PE(n) in the figure. The GPU 200 may include one or more caches and / or buffers 206 configured to receive data 202 from memory 201 and provide processed data 203 to memory 201. Memory 201 may include one or more data storage units arranged in any suitable manner. Typically, memory 201 will include dedicated memory for the GPU and system memory supporting the data processing system of the GPU.

[0058] The individual units of GPU 200 can communicate via one or more data buses and / or interconnects 205. The GPU may include firmware 207, for example, to provide low-level control over the individual units of the GPU.

[0059] Each processing element 204 of the GPU is operable to process a task, wherein the processing elements are arranged such that multiple processing elements can each execute a corresponding task simultaneously. In this way, the GPU can process multiple tasks concurrently. Each processing element may include multiple configurable functional elements (e.g., shaders, geometry processors, vector processors, rasterizers, texture units, etc.) to enable a given processing element to be configured to perform a series of different processing actions. A processing element can process a task by performing a set of actions on a portion of the data used for the task. This set of actions can be defined to suit a given task. Processing elements can be configured, for example, by the GPU's software driver passing appropriate commands to firmware 207 to enable / disable the functional elements of the processing element so that the processing element performs different sets of processing actions. In this way, a first set of processing elements can be configured, for example, to perform vector processing on sensor data received from vehicle sensors, while another set of processing elements can be configured, for example, to perform shader processing on a graphics task representing a portion of a computer-generated image of a scene (e.g., tiles). Each processing element may be able to process a task independently of any other processing element. Therefore, a task processed at one processing element may not cooperate with another processing element to process the task (for example, a single task may not be processed in parallel at more than one processing element, but a single task may be processed in parallel at a single processing element).

[0060] When processing a task, processing element 204 generates output for that task. The output data can be the final output used to process the task, or intermediate output data generated during the processing of the task. GPU 200 includes a checking unit 208 operable to receive the output data from the processing element and form a signature that is characteristic of the output data. For example, the signature can be a characteristic of the output data as it is output from the processing element. In other words, the signature can be a characteristic of the output data as it is output from the processing element. The checking unit can determine, for example, a checksum, hash, cyclic redundancy check (CRC), or fingerprint calculation for the output data. The checking unit can operate on data generated by the processing element processing the task. This data can include memory addresses and / or control data associated with the generated data—this can help the verification operations described herein identify a wide range of faults. The signature provides an expression of the processing performed on the task by the processing element in a more compact form than the output data itself, in order to facilitate comparison of output data provided by different processing elements. Preferably, the inspection unit forms a signature on all output data about the task received from the processing element (which may exclude any control data), but the signature may be formed on some (e.g., not all) of the output data about the task received from the processing element. The inspection unit 208 can receive output data from the processing element via the data bus / interconnect 205.

[0061] The inspection unit 208 may include a data storage area 209 for storing one or more signatures formed at the inspection unit. Alternatively or additionally, the inspection unit may utilize a data storage area outside the inspection unit (e.g., at the memory of GPU 200) to store one or more signatures formed at the inspection unit. The inspection unit may receive output data from all or a subset of the processing elements of the GPU. The inspection unit may include multiple inspection unit instances, for example, each inspection unit instance may be configured to receive output data from a different subset of the processing elements of the GPU.

[0062] GPU 200 further includes a fault detection unit 210 configured to compare two or more signatures formed at inspection unit 208. Fault detection unit 210 is configured to issue a fault signal 211 when a signature mismatch is determined. A fault may potentially lead to a security breach at the GPU. The fault signal can be provided as an output of GPU 200 in any suitable manner. For example, the fault signal can be one or more of the following: control data; an interrupt; data written to memory 201; and data written to registers or memory of GPU 200 or to a system to which GPU is connected.

[0063] Fault detection unit 210 is used to compare signatures of output data from different processing elements 204 arranged to process the same task. The task may be processed multiple times (e.g., twice) by one or more processing elements. The processing performed by the processing elements(s) for multiple processing tasks may or may not be simultaneous. If two processing elements are arranged to process the same task, the signatures of characteristics of the output data output from the processing elements are compared to indicate whether the processing performed by that pair of processing elements is consistent. When the signatures of a pair of processing elements for a given task do not match, fault signal 211 indicates a fault has occurred at one of the processing elements in the pair, but the fault signal does not indicate which processing element experienced the fault.

[0064] If a task is processed three or more times (e.g., by a group of three or more processing elements arranged to process the task), the signatures of the output data from the processing element processing the task are compared to indicate whether the processing performed by that processing element is consistent. In this example, when three or more signatures determined based on the processing of the task do not match, fault signal 211 indicates a fault has occurred at one of the processing elements, and it may further indicate at which processing element the fault occurred. This is because it can be assumed that a fault has occurred at a processing element whose signature does not match the signatures of the outputs from two or more other processing elements.

[0065] GPU 200 can be combined with, for example Figure 3The data processing system 300 shown is a data processing system. Such a data processing system may include other processors (e.g., a central processing unit (CPU) 304) and memory 201. Hardware 302 may include one or more data buses and / or interconnects 308 through which processors 200, 304 and memory 201 can communicate. Typically, a software environment 301 is provided at the data processing system, in which multiple processes 307 can execute. An operating system 306 may provide an abstraction of the available hardware 302 to the processes 307. The operating system may include a driver 309 for GPU 200 to expose the functionality of GPU 200 to the processes. All or part of the software environment 301 may be provided as firmware. In the example, the data processing system 300 forms part of a vehicle control system, wherein each of the processes performs one or more control functions of the vehicle, such as instrument panel display, entertainment system, engine management, air conditioning control, lane control, steering correction, automatic braking system, etc. One or more of the processes 307 may be safety-critical processes. The process can be a mixture of security-critical processes that must be executed according to a predefined security level and non-security-critical processes that do not need to be executed according to a predefined security level.

[0066] The fault signal can be used in any way by the data processing system 300 that incorporates the GPU. For example, when a fault signal is issued by the fault detection unit, the system incorporating the GPU can discard output data about the subject task and / or resubmit the task to the GPU for reprocessing. The GPU itself can use the fault signal 211. For example, the GPU can record fault signals and processing elements associated with these faults, and if one or more processing elements exceed a predefined number of faults (possibly within a limited time period), these one or more processing elements may be disabled or otherwise prevented from processing the task received at the GPU.

[0067] Figure 2 The GPU shown is operable to process tasks in order to meet predefined safety levels. For example, the graphics processing system can be certified to meet the ASIL B or ASIL D standard of ISO 26262. Tasks that need to be processed to the predefined safety level can be tasks related to safety-critical functions of the data processing system 300 in which the GPU can be incorporated. For example, in automotive applications, safety-critical tasks could be those related to image processing of data captured by one or more vehicle cameras used in lane assist systems.

[0068] As described in this article, the tasks to be processed at the data processing system can be test tasks. These can be performed within the processing units of the processing system (e.g., ...). Figure 2The test task is received at the GPU 200 shown. The test task can be used to verify the processing logic of the processing unit. The test task includes a predefined set of instructions for execution on the processing logic (e.g., a processing element). The predefined set of instructions is configured to perform a predetermined set of operations on the processing logic when executed against predefined input data. For example, the test task might need to perform a specific set of data manipulation operations to target a specific set of logic on the processing element, or it might specify a set of reads / writes to be performed to target certain paths to / from memory. The predefined set of instructions can be configured to perform different predetermined sets of operations on the processing logic when executed against different predefined input data. The test task verifies that it is programmed to use a hardware arrangement (e.g., ...). Figure 2 The diagram shows a subset of logic on the GPU 200. In other words, a test task verifies that a specific set of logic it is programmed to use is functioning correctly. In other words, a test task represents a method of hardware testing that involves providing a component with a specific and specially designed task as a stimulus to see if the component delivers the expected results. Different test tasks can be designed for different hardware arrangements (e.g., different graphics processing units). That is, the specific predefined set of instructions defining a test task can vary depending on the hardware arrangement and the capabilities of the processing logic to be verified. Technicians (e.g., software engineers) will be able to design appropriate test tasks on the instructions suitable for the processing logic to be verified, based on the principles described herein.

[0069] Reference Figure 4 and Figure 5 This document describes an exemplary method for processing a task (referred to herein as an input task) at a data processing system based on the principles described herein. It should be understood that the input task can be performed by a process (e.g., ...) executed at the processing system. Figure 3 The process (one of the multiple processes 307 shown) is generated, and can be generated by another task processed at the processing system (e.g., proliferating), or can be generated in any other suitable manner. In short, the input task is copied to form a first task and a second task. Memory to be accessed during the processing of the first and second tasks is allocated, as will be described in further detail herein. The first and second tasks are processed by the processing logic of the processing system to generate a first output and a second output. For example, the first task is processed at a first processing element, and the second task is processed at a second processing element to generate the first output and the second output, respectively. A first signature and a second signature can be formed, the first signature and the second signature being characteristics of the first output and the second output, respectively. The first signature and the second signature can be compared, and if the first signature and the second signature do not match, a fault signal can be issued.

[0070] Figure 4 It shows Figure 3The diagram illustrates an exemplary logical arrangement of units in a data processing system, which are used to process tasks in this manner. Figure 4 A first processing element 204a and a second processing element 204b are shown. The first processing element 204a and the second processing element 204b may have the same characteristics as the reference. Figure 2 The processing element 204 described has the same properties. Figure 4 A first inspection unit 208a and a second inspection unit 208b are also shown. The first inspection unit 208a can be configured to inspect the output generated by the first processing element 204a. The second inspection unit 208b can be configured to inspect the output generated by the second processing element 204b. The first inspection unit 208a and the second inspection unit 208b can be referenced. Figure 2 An example of the described inspection unit 208. Figure 4 A first filter unit 400a and a second filter unit 400b are shown, as will be described in further detail herein. For ease of explanation, the first filter 400a and the second filter 400b are... Figure 4 The first filter 400a and the second filter 400b are shown as logically separate units from the first inspection unit 208a and the second inspection unit 208b; however, the first filter 400a and the second filter 400b can actually be... Figure 2 This is a portion of the inspection unit 208 shown. The first filter 400a and the second filter 400b can be implemented in hardware (e.g., fixed-function circuitry), software, or any combination thereof. The first filter unit 400a can be configured to filter the output generated by the first processing element 204a. The second filter unit 400b can be configured to filter the output generated by the second processing element 204b. The outputs of the first filter unit 400a and the second filter unit 400b can be received at memory level 402 of the data processing system. Figure 4The memory hierarchy 402 shown includes a first L0 cache 206-0a and a second L0 cache 206-0b. The first L0 cache 206-0a is accessible by a first processing element 204a. That is, the first processing element 204a can output instructions requesting to read data from or write data to the first L0 cache 206-0a, while the second processing element 204b may not. The second L0 cache 206-0b is accessible by a second processing element 204b. That is, the second processing element 204b can output instructions requesting to read data from or write data to the second L0 cache 206-0b, while the first processing element 204a may not. The first L0 cache 206-0a and the second L0 cache 206-0b can be local to the GPU that includes the first processing element 204a and the second processing element 204b (e.g., implemented as a GPU on the same physical chip). The first L0 cache 206-0a and the second L0 cache 206-0b can be filled with data from the L1 cache 206-1. The L1 cache 206-1 can be accessed by both the first processing element 204a and the second processing element 204b. That is, both the first processing element 204a and the second processing element 204b can output instructions requesting to read data from or write data to the L1 cache 206-1. The L1 cache 206-1 can be local to the GPU including the first processing element 204a and the second processing element 204b (e.g., implemented as a GPU on the same physical chip). The L1 cache 206-1 can be filled with data from memory 201 (e.g., having a... Figure 2 Data filling (or data of the same nature as memory 201 shown in Figure 3). Memory 201 may not be local to the GPU including the first processing element 204a and the second processing element 204b (e.g., implemented on the same physical chip). Memory hierarchy 402 may include one or more additional cache hierarchies (e.g., L2 cache, or L2 and L3 caches, etc.) between L1 cache 206-1 and memory 201. Figure 4 (Not shown in the image). Figure 4 The processing system shown also includes a task copying unit 404, which is configured to copy input tasks to form a first task and a second task, as will be described in further detail herein. Figure 4 The processing system shown also includes a memory allocation unit 406, which is configured to allocate memory to be accessed during the processing of the first and second tasks, as will be described in further detail herein. The task copying unit 404 and the memory allocation unit 406 can be configured in the processing system's driver (e.g., ...). Figure 3The task copying unit 404 and the memory allocation unit 406 are implemented at the driver 309 shown. They can be implemented in hardware, software, or any suitable combination thereof.

[0071] Figure 5 This document illustrates a method for processing a task at a data processing system based on the principles described herein. The task may be a security task to be processed according to a predefined security level. The task may be a computational workgroup comprising one or more computational work items as described herein. The task may be a test task as described herein.

[0072] In step S502, the input task is copied to form a first task and a second task. For example, the first task may be referred to as a "mission" task, and the second task may be referred to as a "safety task" or a "redundant task." This task can be copied by the task copying unit 404. In an example, copying the input task may include creating a copy of the task. For example, the second task may be defined by a copy of each instruction or line of code defining the first task. In another example, copying the input task may include invoking the input task twice (e.g., without creating a copy of the input task). That is, the input task may be defined by a program stored in memory (e.g., memory 201). The input task can be invoked for processing by providing a processing element that references the program in memory. Therefore, the input task can be copied by the task copying unit 404, which provides a reference to the memory to the processing element that will process the first task, and provides the same reference to the memory to the processing element that will process the second task.

[0073] In step S504, memory to be accessed during the processing of the first and second tasks is allocated. That is, one or more portions of memory 201 are allocated to store data to be accessed during the processing of the first and second tasks. The memory may be allocated by memory allocation unit 406.

[0074] Different types of data can be accessed during task processing. One example is "read-only" data, which is data that the processing element is allowed to read but not write to. That is, the processing element is not allowed to write to memory addresses that contain read-only data. Another type of data that can be accessed during task processing is "read-write" data, which is data that the processing element is allowed to read, modify, and write back to memory. In other words, the processing element can read read-write data from memory, modify the data according to the task being processed, and write the modified data back to memory.

[0075] According to the principles described herein, the allocated memory to be accessed during the processing of a first task and a second task includes a first memory block, a second memory block, and a third memory block. The first memory block is configured to store read / write data to be accessed during the processing of the first task, the second memory block is configured to store a copy of that read / write data to be accessed during the processing of the second task, and the third memory block is configured to store read-only data to be accessed during both the processing of the first and second tasks. The first and second memory blocks may be referred to as "read / write" buffers. The third memory block may be referred to as a "read-only" buffer. During the processing of the second task, the first memory block may not be accessed. That is, the processing element processing the second task may not modify (e.g., write modified read / write data) the first memory block during the processing of the second task. During the processing of the first task, the second memory block may not be accessed. That is, the processing element processing the first task may not modify (e.g., write modified read / write data) the second memory block during the processing of the first task.

[0076] The memory allocation unit 406 allocates the first memory block and the second memory block in such a way that the first processing element and the second processing element do not share access to the same instance of read / write data. Instead, the first processing element is allowed to access the read / write data stored in the first memory block, while the second processing element is allowed to access a copy (e.g., a duplicate) of the read / write data in the second memory block. The reason for allocating the first and second memory blocks in this way is that if, for example, the first and second processing elements are allowed to share access to the read / write data during the processing of the first and second tasks, and the first processing element processing the first task reads the data and executes a set of instructions that modify the data before writing back the modified data, and then the second processing element processing the second task attempts to access the original read / write data to execute the same set of instructions, then the second task will actually access the modified read / write data, and therefore executing the same set of instructions will result in different outputs. If this occurs, then the checking unit 208 (e.g., via the first checking unit instance 208a and the second checking unit instance 208b) will identify the mismatch in the outputs of the first and second processing elements and thereby issue a fault signal, even if the first and second processing elements themselves are operating normally.

[0077] In contrast, since the first and second processing elements are not allowed to modify or write read-only data, they can be allowed to share access to the read-only data. That is, because the read-only data cannot be modified by either processing element, it can be ensured that two processing elements accessing a shared memory address configured to store read-only data will access the same read-only data, even if one processing element accesses the data after the other. Therefore, memory allocation unit 406 can allocate a third memory block configured to store the read-only data to be accessed during the processing of both the first and second tasks.

[0078] refer to Figure 6a , Figure 6b and Figure 6c A more detailed description of memory allocation. Figure 6a A memory stack 600 is shown, in which a first memory block 606, a second memory block 608, and a third memory block 610 have been allocated. The memory stack can be a contiguous block of memory reserved for storing data to process one or more input tasks at the data processing system. The memory stack can be located in the memory 201 of the data processing system. Figure 6a A first memory block 606 is shown configured to store read / write data to be accessed during processing a first task, and a second memory block 608 is configured to store a copy of that read / write data to be accessed during processing a second task. The first memory block 606 may span a first range of contiguous memory addresses in a memory stack 600. The second memory block 608 may span a second range of contiguous memory addresses in a memory stack 600. Each memory address of the first memory block 606 may be offset by a fixed memory address step 612 from a corresponding memory address in the second memory block 608. For example, the base (e.g., first) memory address of the first memory block 606 may be offset by a fixed memory address step 612 from a corresponding base (e.g., first) memory address in the second memory block 608. Corresponding memory addresses in the first and second memory blocks may be configured to store the "same" read / write data. That is, read / write data may be stored at a memory address in the first memory block 606, while a copy of that read / write data may be stored at a corresponding memory address in the second memory block 608.

[0079] It should be understood that mapping corresponding memory addresses in the first and second memory blocks to each other using a fixed memory address step is given by way of example only, and other methods can be used for mapping corresponding memory addresses in the first and second memory blocks. For example, a lookup table can be used to map corresponding memory addresses in the first and second memory blocks to each other. The lookup table can be stored in memory 201. Memory allocation unit 406 can be responsible for filling the lookup table with the mapping between corresponding memory addresses in the first and second memory blocks. In this example, no fixed relationship is required between corresponding memory addresses within the first and second memory blocks.

[0080] Figure 6a A third memory block 610 is also shown, which is configured to store read-only data to be accessed during the processing of both the first and second tasks. The third memory block 610 may span a third range of contiguous memory addresses within the memory stack 600.

[0081] Multiple input tasks can be processed at the data processing system, and for each corresponding copy pair of the first and second tasks, the fixed memory address step size can be the same. This can be referenced. Figure 6b To understand this, the memory heap 600 can be conveniently "divided" into a first sub-heap 602 and a second sub-heap 604. The first sub-heap 602 can be referred to as the "mission heap," and the second sub-heap 604 can be referred to as the "safety sub-heap." The base (e.g., first) memory address of the first sub-heap 602 can be offset from the corresponding base (e.g., first) memory address of the second sub-heap 604 by a fixed memory address step 612. In this example, the fixed memory address step 612 can be conveniently defined as half the size of the memory heap 600. In this way, when a first memory block to be accessed during the processing of a mission task is allocated within the mission heap 602, the corresponding second memory block to be accessed during the processing of a corresponding safety task will be allocated within the safety heap 604 using this fixed memory address step 612. For example, Figure 6b A memory stack 600 is shown, in which a first memory block 606a, a second memory block 608a, and a third memory block 610a have been allocated for access during processing of a first task and a second task associated with a first input task, and a first memory block 606b, a second memory block 608b, and a third memory block 610b have been allocated for access during processing of a first task and a second task associated with a second input task.

[0082] For reference Figure 6bAllocating memory as described is convenient because managing memory resource allocation is computationally simpler when using the same fixed memory address step size of 612 for each input task. Additionally, this makes it easier for drivers (e.g., Figure 3 The driver 309 in the middle can expose only half of the available memory (e.g., indicating its availability) to the process that issues the input task (e.g., Figure 3 The process 307 manages memory resources. For example, if 12GB of memory is actually available, the driver 309 can indicate to the process 307 that 6GB of memory is available during input task processing. The process 307 is unaware that the driver 309 is configured to copy the read and write data to be accessed during input task processing, and therefore such copying will not be considered in the amount of memory it requests. By exposing only half of the available memory to the process 307, it is ensured that there is enough memory available (as requested by the process 307) to store the data to be accessed during the processing task, as well as copies of the read and write data. This is achieved by using the same fixed memory address step size 612 and always as Figure 6b The allocation of memory in mission heap 602 for read-only buffers 610a and 610b, as shown, also ensures that if space is available for read-write buffers 606a and 606b in mission heap 602, the corresponding memory block will always be available for a copy of that read-write buffer 608a and 608b in secure heap 604. However, a potential drawback of this approach is that secure heap 604 may become sparsely populated.

[0083] Therefore, alternatively, for multiple input tasks, the fixed memory address step size can be variable between corresponding copy pairs of the first and second tasks. In other words, for each individual pair of the first and second tasks formed by a specific input task, the memory address step size can be "fixed," but the memory address step size applied to different pairs of the first and second tasks formed by different corresponding input tasks can be different. This can be referenced. Figure 6c To understand. Figure 6cA memory heap 600 is shown, in which a first memory block 606a, a second memory block 608a, and a third memory block 610a have been allocated for access during the processing of a first task and a second task associated with a first input task, and a first memory block 606b, a second memory block 608b, and a third memory block 610b have been allocated for access during the processing of a first task and a second task associated with a second input task. Here, the fixed memory address step 612a between the first memory block 606a and the second memory block 608a of the first input task is greater than the fixed memory address step 612b between the first memory block 606b and the second memory block 608b of the second input task. The fixed memory address step for each input task can be dynamically determined by the driver based on available memory. Additionally, Figure 6c It is shown that third memory blocks 610a and 610b can be dynamically allocated, for example, in any part of the memory stack 600 based on memory availability. This method is advantageous because, relative to a reference... Figure 6b The described method allows the memory heap to be stacked more efficiently (e.g., less sparsely). Even so, allocating memory in this way may be computationally less manageable.

[0084] exist Figure 6b and Figure 6c The diagram illustrates that a first input task and a second input task are each allocated a third memory block, which is configured to store read-only data to be accessed during each corresponding copy pair of processing the first and second tasks (e.g., boxes 610a and 610b). It should be understood that this is not necessarily the case—(e.g., where the first and second input tasks reference the same read-only data) the first and second input tasks may share access to a single third memory block configured to store read-only data to be accessed during each corresponding copy pair of processing the first and second tasks.

[0085] Return to Figure 5 In step S506, a first task is processed at the first processing element 204a and a second task is processed at the second processing element 204b to generate a first output and a second output, respectively. The first processing element 204a and the second processing element 204b can be identical. That is, the first processing element 204a and the second processing element 204b can include the same hardware and be configured in the same way (e.g., by...). Figure 2 The firmware 207 configuration may be in the system. Figure 3(As instructed by driver 307 in the document). Moreover, as described herein, the first task and the second task are copies. Therefore, in the absence of any faults, the calculations occurring at each of the first processing element 204a and the second processing element 204b during the processing of the first task and the second task respectively should be identical, thus producing matching outputs.

[0086] Drivers (e.g., Figure 3 The driver (309) can be responsible for submitting (e.g., scheduling) the first and second tasks for processing on the first processing element 204a and the second processing element 204b, respectively. The first and second tasks can be submitted by the driver in parallel (e.g., simultaneously) to the first processing element 204a and the second processing element 204b, respectively. As described above, the L1 cache 206-1 can be filled with data from memory 201 (e.g., using data from memory allocated in step S504), while the first L0 cache 206-0a and the second L0 cache 206-0b can be filled with data from the L1 cache 206-1. Therefore, by submitting the first and second tasks in parallel, the data processing system may benefit (e.g., in the L1 cache 206-1) from the data that the cache will access during the processing of these tasks. In other words, by submitting the first task and the second task simultaneously, it is likely that the data (or at least a large portion thereof) to be accessed during the processing of these tasks can be retrieved from memory 201 at once and cached in L1 cache 206-1 so that it can be accessed by both the first task and the second task, instead of each processing element having to retrieve the data from memory 201 separately.

[0087] Although the first task and the second task can be submitted to the first processing element 204a and the second processing element 204b simultaneously, the first processing element 204a and the second processing element 204b can be configured to process the first task and the second task independently, respectively. That is, the first and second processing elements 204a do not need to synchronize for each copy pair of the first task and the second task in order to execute these tasks simultaneously on a cycle-by-cycle basis.

[0088] As described herein, a first task is processed at a first processing element 204a and a second task is processed at a second processing element 204b to generate a first output and a second output, respectively. The output can be the final output of processing the task or an intermediate output generated during the processing of the task. For example, an intermediate output can be one or more of load, store, or atomic instructions generated during the processing of the task. An intermediate output can include a reference to a memory address. That is, an intermediate output can include a request to access data in memory to be used during the processing of the task.

[0089] Reference Figure 7 Describes a set of exemplary steps performed by each inspection unit 208a, 208b and filter unit 400a, 400b in response to intermediate outputs generated during the processing of the first task and the second task at the first processing element 204a and the second processing element 204b.

[0090] In step S702, the output generated during task processing at the processing element is received at the inspection unit (e.g., inspection unit 208a or 208b). The steps performed at inspection units 208a and 208b are identical, regardless of whether their associated processing element is processing a first (e.g., "mission") task or a second (e.g., "security") task. As described herein, the first and second tasks are duplicates, and the first processing element 204a and the second processing element 204b are preferably identical; therefore, in the absence of any faults, the processing of the first and second tasks should be identical, and thus should produce matching outputs.

[0091] In step S704, the checking unit forms a signature that is a characteristic of the received output. For example, the signature may be a characteristic of the output data as it is output from the processing element. In other words, the signature may be a characteristic of the output data as it is output from the processing element. As described herein, forming a signature that is a characteristic of the received output may include performing one or more of checksum, CRC, hash, and fingerprint on the output, respectively. Preferably, the checking unit forms a signature on all the output data received from the processing element regarding the task (which may include any referenced memory addresses), but may form a signature on some (e.g., not all) of the output data received from the processing element regarding the task. For example, when the input task is copied to form a first task and a second task, (e.g., in...) Figure 3 The task copying unit 404 (implemented at driver 309) may have flagged any references to memory addresses in the first, second, and third memory blocks. The first and second tasks may also include references to other memory addresses not flagged by the task copying unit 404 (e.g., references to memory addresses in overflow registers), which may be non-deterministic. The checking unit may form a signature on the output data, which includes only the flagged memory addresses present in the output data. In this way, non-deterministic references to memory addresses can be excluded from the signature. The signature can be stored (e.g., in...). Figure 2 In the data storage area 209 shown, a signature is compared with a signature formed for the corresponding output generated during the processing of another task in the first and second tasks. The comparison of signatures and the issuance of fault signals, where applicable, have been previously described herein.

[0092] It should be noted that the signature of the characteristics of the received output is preferably formed by the checking unit (e.g., checking unit 208a or 208b) before the output accesses memory level 402. That is, the signature of the characteristics of the received output is preferably formed by the checking unit (e.g., checking unit 208a or 208b) before the output accesses the corresponding L0 cache (e.g., L0 cache 206-0a or 206-0b). This is because the output of the L0 cache may not be decisive, depending on when the cache line is evicted, and therefore the output may differ even if the inputs to the L0 cache are the same. After the signature is formed, the checking unit (e.g., checking unit 208a or 208b) can forward the intermediate output to the corresponding filtering unit (e.g., filtering unit 400a or 400b).

[0093] In step S706, the filtering unit (e.g., filtering unit 400a or 400b) can determine whether the intermediate output includes a reference to a memory address configured to store read / write data. In one example, when copying an input task to form a first task and a second task, (e.g., in...) Figure 3 The task copying unit 404, implemented at driver 309, may have flagged any references to memory addresses configured to store read / write data in the first and second tasks so that the filtering unit can quickly identify these memory addresses.

[0094] If the intermediate output does not include a reference to the memory address configured to store read / write data, it can be forwarded to memory level 402 in step S708.

[0095] If the intermediate output does indeed include a reference to a memory address configured to store read / write data, then in step S710, a filtering unit (e.g., filtering unit 400a or 400b) can determine whether the intermediate output was generated during the processing of the second task. That is, for example, the filtering unit (e.g., filtering unit 400a or 400b) can determine whether the intermediate output was received from the second processing element 204b. In one example, when the second task is submitted for processing by the second processing element, the driver (e.g., ...) Figure 3 The driver (309) can identify the second filter unit 400b as associated with the second processing element 204b, for example, by setting a configuration bit in a register associated with the filter unit 400b. Therefore, the filter unit (e.g., filter unit 400a or 400b) can determine whether the intermediate output was generated during the processing of the second task by checking the configuration bit.

[0096] If it is determined that the intermediate output was not generated during the processing of the second task (e.g., it was generated during the processing of the first task), it can be forwarded to memory level 402 in step S712. At this point, the intermediate output can access read / write data at the referenced memory address (e.g., in the first memory block).

[0097] If it is determined that the intermediate output was generated during the processing of the second task, the reference to the memory address configured to store read / write data can be updated in step S714. That is, the reference to the memory address in the first memory block can be modified by the filtering unit 400b to reference the corresponding memory address in the second memory block. In this example, this can depend on a fixed memory address stride, for example, by adding a fixed memory address stride to the memory address in the first memory block to determine the corresponding memory address in the second memory block. In another example, this can be achieved by referencing a lookup table to map the memory address in the first memory block to the corresponding memory address in the second memory block. The updated intermediate output can be forwarded to memory level 402, where the corresponding memory address in the second memory block can be accessed at the updated referenced memory address. For example, accessing the corresponding memory address in the second memory block can include reading read / write data from the second memory block and returning that data to the second processing element. In this example, the data returned from memory level 402 may not include a reference to the memory address and therefore does not need to be routed via the filtering unit 400b. In another example, accessing a corresponding memory address in a second memory block may include writing read / write data modified by a second processing element to the corresponding memory address.

[0098] That is, when determining the intermediate output: (i) includes a reference to a memory address configured to store read / write data, and (ii) is generated during the processing of the second task, the filtering unit (e.g., filtering unit 400a or 400b) can update the intermediate output. Therefore, it should be understood that steps S706 and S710 can be referenced... Figure 7 The order described is the reverse order of execution.

[0099] It should be noted that, in reference Figure 7In the described example, filtering units 400a and 400b are logically positioned after checking units 208a and 208b. In this way, updating the read / write memory address referenced in the output by the second processing element does not cause inconsistencies in the signature formed by checking units 208a and 208b. It should be understood that this is not necessarily the case. For example, checking units 208a and 208b can be configured to form a signature that does not consider memory addresses referenced in intermediate outputs. In this example, filtering units 400a and 400b can be logically positioned before checking units 208a and 208b, respectively. Alternatively, in this example, the driver (e.g., Figure 3 The driver (309) can update the memory address in the first memory block in the second task before submitting any reference to it for processing at the second processing element 204b. In this case, the filtering units 400a and 400b are not required.

[0100] If the processing of the first and second tasks is completed and no fault signal is issued in response to any intermediate or final output (e.g., all corresponding intermediate and final outputs match), then the final processed output of the first or second task can be regarded as the processed output of the input task.

[0101] like Figure 4 As shown, the first task and the second task are processed by a first processing element and a second processing element, which are different processing elements. It should also be understood that in other examples according to the principles described herein, the first task and the second task may be processed by the same processing element, for example, at different times. That is, the first task and the second task may be submitted for processing by the same processing element. In the first round, the processing element may process the first task. A checking unit associated with the processing element may form a signature for the output generated during processing the first task and store these signatures. In the second round, the processing element may process the second task. The checking unit may form a signature for the output generated during processing the second task and compare it with the corresponding stored signature formed for the first task. A filtering unit associated with the processing element may be able to determine whether the processing element is performing the first round or the second round and update the reference to a memory address configured to store read / write data in the intermediate output generated during processing the second task. For example, configuration bits in a register associated with the processing element may be used to indicate whether the processing element is performing the first round (e.g., binary "0") or the second round (e.g., binary "1").

[0102] In another example, Figure 4The data processing system shown may further include a third processing element, a third inspection unit, a third filtering unit, and a third L0 cache. (For example, in...) Figure 3 The task copying unit 404, implemented at driver 309, can copy the input task to form a third task. The memory allocation unit 406 can allocate additional memory including a fourth memory block configured to store copies of read / write data accessed during the processing of the third task. The principles described herein can be applied to processing the third task at a third processing element. In this example, the output of the input task can still be provided even if one of the first, second, or third processing elements experiences a failure. That is, the output of the input task can be provided as the output of two matching processing elements, while the other processing element providing a mismatched output can be considered faulty.

[0103] In yet another example, the allocated memory to be accessed during processing the first and second tasks may further include a fourth memory block configured to store write-only data generated during processing the first task and a fifth memory block configured to store corresponding write-only data generated during processing the second task. Alternatively, the allocated memory to be accessed during processing the first and second tasks may further include a fourth memory block configured to store only write-only data generated during processing the first task. In this example, the filtering unit 400b may be configured to filter out (e.g., block) writes to write-only data by the second processing element processing the second (e.g., "safe" or "redundant") task in order to reduce latency associated with processing the second task and / or save bandwidth. In this example, the final output generated during processing the first task may be used as the output for processing the input task (assuming no fault signal is issued during processing the first and second tasks).

[0104] Figures 2 to 4 The data processing system is shown as comprising multiple functional blocks. This is merely illustrative and not intended to define a strict division between the different logical elements of such an entity. Each functional block may be provided in any suitable manner. It should be understood that the intermediate values ​​described herein formed by the data processing system do not need to be physically generated by the data processing system at any point in time, and may merely represent logical values ​​that conveniently describe the processing performed by the data processing system between its inputs and outputs.

[0105] The data processing system described herein may be contained in hardware on an integrated circuit. The data processing system described herein may be configured to perform any of the methods described herein. Generally, any of the functions, methods, techniques, or components described above may be implemented in software, firmware, hardware (e.g., a fixed logic circuit system), or any combination thereof. The terms “module,” “function,” “component,” “element,” “cell,” “block,” and “logic” may be used herein to generally denote software, firmware, hardware, or any combination thereof. In the case of a software implementation, a module, function, component, element, cell, block, or logic represents program code that, when executed on a processor, performs a specified task. The algorithms and methods described herein may be executed by one or more processors executing code that causes the processor to execute the algorithm / method. Examples of computer-readable storage media include random access memory (RAM), read-only memory (ROM), optical disk, flash memory, hard disk storage, and other memory devices that may use magnetic, optical, and other techniques to store instructions or other data and may be accessible by a machine.

[0106] As used herein, the terms computer program code and computer-readable instructions refer to any kind of executable code for a processor, comprising code expressed in machine language, interpreted language, or scripting language. Executable code includes binary code, machine code, bytecode, code defining integrated circuits (e.g., hardware description languages ​​or netlists), and code expressed in programming languages ​​such as C, Java, or OpenCL. Executable code can be, for example, any kind of software, firmware, script, module, or library that, when properly executed, processed, interpreted, compiled, or run in a virtual machine or other software environment, causes the processor of a computer system that supports the executable code to perform tasks specified by said code.

[0107] A processor, computer, or computer system can be any kind of device, machine, or special-purpose circuit, or a collection or part thereof, that has the processing capability to execute instructions. A processor can be or includes any kind of general-purpose or special-purpose processor, such as a CPU, GPU, NNA, system-on-a-chip, state machine, media processor, application-specific integrated circuit (ASIC), programmable logic array, field-programmable gate array (FPGA), etc. A computer or computer system may include one or more processors.

[0108] This invention also intends to cover software, such as hardware description language (HDL) software, that defines the configuration of hardware as described herein, for designing integrated circuits or for configuring programmable chips to perform desired functions. That is, a computer-readable storage medium may be provided on which computer-readable program code in the form of an integrated circuit definition dataset is encoded, which, when processed (i.e., run) in an integrated circuit manufacturing system, configures the system to manufacture a data processing system configured to perform any of the methods described herein, or to manufacture a data processing system including any of the devices described herein. The integrated circuit definition dataset may, for example, be an integrated circuit description.

[0109] Therefore, a method for manufacturing a data processing system as described herein can be provided at an integrated circuit manufacturing system. Furthermore, an integrated circuit definition dataset can be provided, which, when processed in the integrated circuit manufacturing system, enables the method for manufacturing the data processing system to be executed.

[0110] Integrated circuit definition datasets can be in the form of computer code, such as as a netlist, code for configuring programmable chips, or as a hardware description language suitable for manufacturing at any level in integrated circuits, including as register-transfer level (RTL) code, as high-level circuit representations (such as Verilog or VHDL), and as low-level circuit representations (e.g., OASIS(RTM) and GDSII). Higher-level representations (such as RTL) that logically define hardware suitable for manufacturing in integrated circuits can be processed on a computer system configured to generate manufacturing definitions of integrated circuits within the context of a software environment that includes definitions of circuit elements and rules for combining these elements to generate the manufacturing definitions of integrated circuits defined by the representation. As is typically the case where software executes at a computer system to define a machine, one or more intermediate user steps (e.g., providing commands, variables, etc.) may be required to configure the computer system to generate manufacturing definitions of integrated circuits, executing code that defines the integrated circuits to generate the manufacturing definitions of the integrated circuits.

[0111] Now about Figure 8 Describe an example of processing integrated circuit definition datasets at an integrated circuit manufacturing system in order to configure the system as a manufacturing data processing system.

[0112] Figure 8An example of an integrated circuit (IC) manufacturing system 802 is shown, configured to manufacture a data processing system as described in any of the examples herein. Specifically, the IC manufacturing system 802 includes a layout processing system 804 and an integrated circuit generation system 806. The IC manufacturing system 802 is configured to receive an IC definition dataset (e.g., defining a data processing system as described in any of the examples herein), process the IC definition dataset, and generate an IC (e.g., containing the data processing system as described in any of the examples herein) based on the IC definition dataset. The processing of the IC definition dataset configures the IC manufacturing system 802 to manufacture an integrated circuit containing the data processing system as described in any of the examples herein.

[0113] The layout processing system 804 is configured to receive and process an IC definition dataset to determine a circuit layout. Methods for determining a circuit layout based on an IC definition dataset are known in the art and may involve, for example, synthesizing RTL code to determine the gate-level representation of the circuit to be generated, for example, in relation to logic components (e.g., NAND, NOR, AND, OR, MUX, and FLIP-FLOP components). By determining the location information of the logic components, the circuit layout can be determined based on the gate-level representation of the circuit. This can be done automatically or with user intervention to optimize the circuit layout. When the layout processing system 804 has determined the circuit layout, it can output the circuit layout definition to the IC generation system 806. The circuit layout definition may be, for example, a circuit layout description.

[0114] IC generation system 806 generates ICs according to circuit layout definitions as known in the art. For example, IC generation system 806 can implement a semiconductor device manufacturing process for generating ICs, which may include a multi-step sequence of photolithography and chemical processing steps, during which electronic circuits are progressively formed on a wafer made of semiconductor material. The circuit layout definition may be in the form of a mask, which can be used in the photolithography process to generate ICs according to the circuit definition. Alternatively, the circuit layout definition provided to IC generation system 806 may be in the form of computer-readable code, which IC generation system 806 can use to form an appropriate mask for generating ICs.

[0115] The different processes performed by the IC manufacturing system 802 can all be implemented in one location, for example, by one party. Alternatively, the IC manufacturing system 802 can be a distributed system, such that some processes can be performed at different locations and by different parties. For example, some of the following levels can be performed at different locations and / or by different parties: (i) synthesizing RTL code representing an IC definition dataset to form a gate-level representation of the circuit to be generated; (ii) generating a circuit layout based on the gate-level representation; (iii) forming a mask based on the circuit layout; and (iv) using the mask to manufacture the integrated circuit.

[0116] In other examples, processing the integrated circuit definition dataset at the integrated circuit manufacturing system can configure the system as a manufacturing data processing system, without processing the IC definition dataset to determine circuit layout. For example, the integrated circuit definition dataset can define the configuration of a reconfigurable processor such as an FPGA, and processing the dataset can configure the IC manufacturing system (e.g., by loading the configuration data into the FPGA) to generate a reconfigurable processor with the defined configuration.

[0117] In some embodiments, when processed in an integrated circuit manufacturing system, the integrated circuit manufacturing definition dataset can enable the integrated circuit manufacturing system to generate devices as described herein. For example, the integrated circuit manufacturing definition dataset referenced above... Figure 8 The configuration of the integrated circuit manufacturing system described herein enables the manufacture of data processing systems as described in this article.

[0118] In some examples, an integrated circuit definition dataset may contain software running on hardware defined at the dataset, or software running in combination with hardware defined at the dataset. Figure 8 In the example shown, the IC generation system can be further configured by the integrated circuit definition dataset to load firmware onto the integrated circuit according to the program code defined in the integrated circuit definition dataset during the manufacturing of the integrated circuit, or otherwise provide the integrated circuit with program code for use with the integrated circuit.

[0119] Compared to known implementations, the implementation of the concepts set forth in this application in devices, apparatuses, modules, and / or systems (and in the methods implemented herein) can lead to performance improvements. Performance improvements may include one or more of increased computational performance, reduced latency, increased throughput, and / or reduced power consumption. During the manufacture of such devices, apparatuses, modules, and systems (e.g., in integrated circuits), trade-offs can be made between performance improvements and physical implementations, thereby improving manufacturing methods. For example, a trade-off can be made between performance improvements and layout area, matching the performance of known implementations but using less silicon. This can be accomplished, for example, by reusing functional blocks serially or sharing functional blocks among elements of a device, apparatus, module, and / or system. Conversely, the concepts set forth in this application that lead to improvements in the physical implementations of devices, apparatuses, modules, and systems (such as reduced silicon area) can be traded off for performance improvements. This can be accomplished, for example, by manufacturing multiple examples of modules within a predefined area budget.

[0120] The applicant has independently disclosed each individual feature described herein, as well as any combination of two or more such features, to the extent that such features or combinations can be implemented based on the specification as a whole, in accordance with the common knowledge of those skilled in the art, regardless of whether such features or combinations of features solve any problem disclosed herein. In view of the foregoing description, those skilled in the art will understand that various modifications can be made within the scope of this invention.

Claims

1. A method for processing input tasks in a processing system, the method comprising: The input task is copied to form a first task and a second task; Allocate memory, the memory comprising: A first memory block, configured to store read / write data to be accessed during the processing of the first task; A second memory block, configured to store a copy of the read / write data to be accessed during the processing of the second task; and A third memory block, configured to store read-only data to be accessed during the processing of both the first and second tasks; and The first task and the second task are processed at the processing logic of the processing system to generate a first output and a second output, respectively.

2. The method according to claim 1, further comprising: A first signature and a second signature are formed, wherein the first signature and the second signature are features of the first output and the second output, respectively. Compare the first signature and the second signature; as well as If the first signature and the second signature do not match, a fault signal is issued.

3. The method of claim 2, further comprising forming the first signature and the second signature before the first output and the second output access the memory level of the processing system.

4. The method according to any one of claims 1 to 3, further comprising, before processing the first task and the second task, storing the read / write data at a memory address of the first memory block, and storing a copy of the read / write data at a corresponding memory address of the second memory block.

5. The method according to any one of claims 1 to 3, wherein, The first memory block and the second memory block are allocated in the memory heap, and each memory address of the second memory block is offset from the corresponding memory address in the first memory block by a fixed memory address step.

6. The method according to claim 5, wherein, Multiple input tasks are processed at the processing system, and the fixed memory address stride is the same for each pair of first and second tasks formed by the corresponding input tasks.

7. The method according to claim 5, wherein, The fixed memory address step size is half the size of the memory heap.

8. The method according to claim 5, wherein, The memory stack is a contiguous block of memory reserved for storing data for processing one or more input tasks at the processing system, and the memory stack is located in the memory of the processing system.

9. The method according to any one of claims 1 to 3, further comprising: Receive the second output; The second output identifies references to memory addresses in the first memory block; Update the citation; as well as Use the updated reference to access the corresponding memory address in the second memory block.

10. The method according to claim 9, wherein, The method further includes allocating a first memory block and a second memory block in a memory heap, wherein each memory address of the second memory block is offset by a fixed memory address step from the corresponding memory address in the first memory block, and updating the reference to the memory address in the first memory block in the second output using the fixed memory address step.

11. The method according to claim 9, further comprising: Receive output, and identify that the output is received from the second task, so as to identify the output as the second output.

12. The method according to any one of claims 1 to 3, further comprising: The first task and the second task are submitted to the processing logic simultaneously.

13. The method according to claim 12, further comprising: Data is extracted from the first memory block, the second memory block, and the third memory block into a cache, which is configured to be accessed by the processing logic during the processing of the first task and the second task.

14. The method according to any one of claims 1 to 3, wherein, The input task is a security task that is processed according to a predefined security level.

15. The method according to any one of claims 1 to 3, wherein, The processing logic includes a first processing element and a second processing element, wherein processing the first task and the second task at the processing logic of the processing system includes processing the first task at the first processing element and processing the second task at the second processing element.

16. The method according to claim 15, wherein, The input task is a test task, which includes a predefined set of instructions for execution on the processing logic. The predefined set of instructions is configured to perform a predetermined set of operations on the processing logic when executed against predefined input data. The method further includes receiving the test task at a processing unit that includes the first processing element and the second processing element.

17. The method according to any one of claims 1 to 3, wherein, The processing logic includes a specific processing element, wherein processing the first task and the second task at the processing logic of the processing system includes processing the first task at the specific processing element and processing the second task at the specific processing element.

18. The method according to any one of claims 1 to 3, wherein, The first output and the second output include intermediate outputs generated during the processing of the first task and the second task, respectively, and optionally, the intermediate outputs are one or more of load, store, or atomic instructions generated during the processing of the task.

19. A processing system configured to process an input task, the processing system comprising: A task copying unit, configured to copy the input task to form a first task and a second task; A memory allocation unit, configured to allocate memory, the memory comprising: A first memory block, configured to store read / write data to be accessed during the processing of the first task; A second memory block, configured to store a copy of the read / write data to be accessed during the processing of the second task; and A third memory block, configured to store read-only data to be accessed during the processing of both the first and second tasks; and The processing logic is configured to process the first task to generate a first output, and to process the second task to generate a second output.

20. The processing system of claim 19 further includes a memory stack, the memory stack comprising the first memory block, the second memory block, and the third memory block.

21. A non-transitory computer-readable storage medium storing a computer-readable description of an integrated circuit, the computer-readable description, when processed in an integrated circuit manufacturing system, causing the integrated circuit manufacturing system to manufacture a processing system configured to process an input task, the processing system comprising: A task copying unit, configured to copy the input task to form a first task and a second task; A memory allocation unit, configured to allocate memory, the memory comprising: A first memory block, configured to store read / write data to be accessed during the processing of the first task; A second memory block, configured to store a copy of the read / write data to be accessed during the processing of the second task; and A third memory block, configured to store read-only data to be accessed during the processing of both the first and second tasks; and The processing logic is configured to process the first task to generate a first output, and to process the second task to generate a second output.

Citation Information

Patent Citations

  • Method and apparatus for fast context cloning in data processing system

    CN110892381A

  • Method and circuit arrangement for synchronization of synchronously or asynchronously clocked processor units

    CN1682195A