System and method for testing components in a system at runtime

By introducing an in-system component tester (ICT) into the SoC or SiP for real-time component testing, the error operation problems caused by minor faults that cannot be discovered after integration are solved, and the high reliability and stability of the SoC or SiP are achieved.

CN114690028BActive Publication Date: 2025-05-30DEEPX CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111170091.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-12-31
Filing Date
2021-10-08
Publication Date
2025-05-30
Estimated Expiration
2041-10-08

AI Technical Summary

Technical Problem

When multiple semiconductor components are integrated into SoC or SiP, the complexity increases significantly, resulting in an increase in defect rate during manufacturing, and minor failures that are not discovered before leaving the factory may gradually amplify during use due to fatigue stress or physical stress, resulting in the wrong operation of SoC or SiP, especially in mission-critical products.

Method used

In-system component tester (ICT) is introduced in SoC or SiP. By monitoring the status of multiple functional components, selecting functional components in an idle state as the components to be tested, performing testing, and interrupting the test steps or deactivating the defective functional components based on detected conflicts or defects, allowing the backup components to be connected to the system bus.

Benefits of technology

It enables component testing during the run time of the SoC or SiP, discovers and deals with potential defects, ensuring high reliability of the system, especially when installed in autonomous vehicles or AI robots.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114690028B_ABST
    Figure CN114690028B_ABST
Patent Text Reader

Abstract

A system-on-chip (SoC) for testing components in a system during runtime includes: a plurality of functional components; a system bus for allowing the plurality of functional components to communicate with each other; one or more wrappers, each wrapper being connected to one of the plurality of functional components; and an in-system component tester (ICT). The ICT monitors the status of the functional components through the wrappers; selects at least one functional component in an idle state as a component under test (CUT); tests the at least one selected functional component through the wrapper; interrupts a test step regarding the at least one selected functional component based on detecting a conflict in accessing the system bus to the at least one selected functional component; and enables the at least one functional component to be connected to the system bus based on the interrupt step.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims priority to Korean Patent Application No. 10 - 2020 - 0189414, filed with the Korean Intellectual Property Office on December 31, 2020, the disclosure of which is incorporated herein by reference. Technical field

[0003] This disclosure relates to testing for faults in a system during operation. Background art

[0004] A system configured by various semiconductor components is implemented through a board - level system based on a printed circuit board (PCB).

[0005] Due to the high integration that can be achieved according to the development of semiconductor manufacturing process technology, a system - on - chip (SoC) or a system - in - package (SiP) is being proposed. In the system - on - chip (SoC), various semiconductor components such as a processor, a memory, and peripherals are implemented in one chip. In the system - in - package (SiP), various semiconductor components such as a processor, a memory, and peripherals are implemented in one package.

[0006] An SoC refers to a semiconductor device (chip) that incorporates an entire system in one chip, and also refers to the technology of implementing main semiconductor components such as arithmetic, memory, and data conversion elements in one chip. A SiP refers to a semiconductor device that incorporates an entire system in one package, and also refers to the technology of implementing main semiconductor components such as arithmetic, memory, and data conversion elements in one package. That is, a central processing unit (CPU) of a computer, a digital signal processing (DSP) chip, a microcontroller unit (MCU), etc. are integrated in one semiconductor die or package, such that the chip or package itself serves as a system. As described above, when semiconductor devices with multiple functions are combined in one chip, the board space is significantly reduced, thereby enabling the reduction of the size of various electronic products. In addition, compared with the technology of manufacturing multiple semiconductor devices separately, the manufacturing cost of semiconductor devices is significantly reduced, thereby also reducing the unit selling price of electronic products. Therefore, the SoC or SiP technology that integrates all component functions is becoming the core component technology in the advanced digital era, which focuses on high performance, low cost, and small size.

[0007] On the other hand, artificial intelligence (AI) is also gradually developing. AI refers to the intelligence that artificially imitates human intelligence, that is, the intelligence for recognition, classification, reasoning, prediction, control / decision, etc. Recently, in order to accelerate the operation speed of artificial intelligence (AI), a neural processing unit (NPU) is being developed. Summary of the invention

[0008] The inventors of the present disclosure have recognized that when an NPU is integrated in a SoC or SiP, the size of the board substrate is reduced, thereby innovatively reducing the size of the electronic device.

[0009] In addition, the inventors of the present disclosure have recognized that when an NPU is integrated in a SoC or SiP, the manufacturing cost can be reduced compared to a semiconductor device manufactured separately.

[0010] However, the inventors of the present disclosure have also recognized that when multiple semiconductor components are integrated in a SoC or SiP, the complexity significantly increases, which increases the defect rate during the manufacturing process. Defects during the manufacturing process may be detected during pre - shipment testing, but some minor defects in components integrated in the SoC or SiP that are not detected during pre - shipment testing may be handed over to the user. Such minor defects gradually amplify due to fatigue stress or physical stress caused by repeated use, ultimately resulting in incorrect operation of the SoC or SiP.

[0011] When the SoC or SiP is installed in an electronic device for user entertainment, its incorrect operation may not be as problematic. However, the inventors of the present disclosure have recognized that the situation is different when the SoC or SiP is installed in a mission - critical product.

[0012] Specifically, the inventors of the present disclosure have recognized the problem that when the NPU in the SoC or SiP malfunctions, is defective, or is damaged, it may output unpredictable artificial intelligence (AI) operation results.

[0013] For example, the inventors of the present disclosure have recognized that when a SoC or SiP including an NPU is used in an electronic device installed in an autonomous vehicle or in an electronic device installed in an AI robot, unpredictable AI operation results may be output due to NPU malfunction, defect, or damage.

[0014] Therefore, the inventors of the present disclosure have recognized the need to propose a method for performing tests in the SoC or SiP at runtime, which has hitherto only been performed before shipment.

[0015] According to one aspect of the present disclosure, there is provided a system-on-chip (SoC) for testing components in a system during runtime. The SoC may include: a plurality of functional components, each of the plurality of functional components including circuitry; a system bus configured to allow the plurality of functional components to communicate with each other; one or more wrappers, each of the one or more wrappers being connected to one of the plurality of functional components; and an in-system component tester (ICT). The ICT is configured to: monitor the states of the plurality of functional components via the one or more wrappers; select at least one of the plurality of functional components that is in an idle state as a component under test (CUT); test the at least one functional component selected as the CUT via the one or more wrappers; interrupt a test step regarding the at least one functional component selected as the CUT based on detecting a conflict with an access from the system bus to the at least one functional component selected as the CUT; and based on the interrupt step, allow the at least one functional component to connect to the system bus.

[0016] The ICT may further be configured to: after allowing the at least one functional component to connect to the system bus, if as a result of the monitoring step, the at least one functional component is again in the idle state, return to the selecting step. The return to the selecting step may occur after a back-off time period regarding the conflict has expired.

[0017] The plurality of functional components includes one or more universal processing units (UPUs). The one or more UPUs may include at least one of the following: one or more central processing units (CPUs); one or more graphics processing units (GPUs); and one or more neural processing units (NPUs) configured to perform operations of an artificial neural network (ANN) model. The plurality of functional components further includes at least one of the following: at least one memory; at least one memory controller; and at least one input and output (I / O) controller.

[0018] For the test step, the ICT may further be configured to instruct the one or more wrappers to isolate the connection of the at least one functional component selected as the CUT from the system bus.

[0019] The ICT may include at least one of the following: a detector configured to monitor the states of a plurality of functional components; a scheduler configured to manage the operations of the ICT; a generator configured to generate test input data; and a tester configured to import the test input data into the CUT and analyze test results obtained from the CUT that processes the test input data. The test input data is predefined test data or a random bit stream generated based on a seed.

[0020] The ICT may further be configured to: after the test step is completed, analyze test results obtained from at least one of the functional components selected as the CUT; and based on the at least one functional component being analyzed as normal, allow the at least one functional component to be connected to the system bus or another system connection.

[0021] The ICT may also be configured to deactivate at least one of the functional components based on the at least one functional component being analyzed as defective. The SoC may further include: a field programmable gate array (FPGA) configured to emulate at least one of the functional components analyzed as defective. The FPGA may have an address that is revoked and replaced by the address of at least one of the functional components analyzed as defective. The deactivation step may include revoking the address of at least one of the functional components analyzed as defective, powering off at least one of the functional components analyzed as defective, and isolating at least one of the functional components analyzed as defective from the system bus by cutting the system bus connection to at least one of the functional components analyzed as defective. The plurality of functional components includes spare components for at least one of the functional components analyzed as defective; and the ICT may also be configured to enable the spare components.

[0022] The test step may be repeatedly performed before and after the SoC is delivered from the factory, and may verify whether there are defects in the manufacturing of the SoC, whether it has been damaged, or whether it has been broken.

[0023] The test step may include a scan test different from a functional test. For the scan test, the ICT may further be configured to: connect a plurality of flip - flops in each CUT to each other; import test inputs into at least one flip - flop; and obtain test results from the operation of the combinational logic of the flip - flops to analyze whether the CUT is defective or normal during runtime.

[0024] The plurality of functional components may include a neural processing unit (NPU). The NPU may include a plurality of arrays of processing elements and may be configured to select and test at least one of the plurality of arrays of processing elements.

[0025] According to another aspect of the present disclosure, there is provided a system-on-chip (SoC) for testing components in a system during runtime. The SoC may include a plurality of functional components that communicate with each other via a system bus; one or more wrappers, each of the one or more wrappers being connected to one of the plurality of functional components; and an in-system component tester (ICT). The ICT is configured to: when it is monitored by the one or more wrappers that at least one of the functional components is in an idle state, select the at least one functional component in the idle state as a component under test (CUT), and based on detecting an access to the selected at least one functional component, allow the selected at least one functional component to be connected to the system bus.

[0026] According to another aspect of the present disclosure, there is provided a method for testing components in a system-on-chip (SoC) during runtime. The method may include: monitoring the states of a plurality of functional components; selecting at least one functional component in the idle state among the plurality of functional components as a component under test (CUT); testing the at least one functional component selected as the CUT; interrupting a test step regarding the at least one functional component selected as the CUT based on detecting a conflict with an access from the system bus to the at least one functional component selected as the CUT; and based on the interrupt step, allowing the at least one functional component to be connected to the system bus.

[0027] According to the present disclosure, tests that were previously only performed before shipment can be performed during runtime in an SoC or SiP.

[0028] According to the present disclosure, it is advantageous that minor faults not detected before shipment in an SoC or SIP are gradually amplified due to fatigue stress or physical stress caused by repeated driving.

[0029] According to the present disclosure, it is advantageous to detect that an NPU in an SoC or SiP outputs unpredictable artificial intelligence (AI) operation results due to a fault, defect, or damage.

[0030] Therefore, according to the present disclosure, high reliability of an SoC or SiP installed in an autonomous vehicle or an AI robot can be ensured.

[0031] Although the present disclosure mainly describes an SoC, the present disclosure is not limited to an SoC, but can be applied to a system-in-package (SIP) or a board-level system based on a printed circuit board (PCB). BRIEF DESCRIPTION OF THE DRAWINGS

[0032] The above and other aspects, features, and other advantages of the present disclosure will be more clearly understood from the following detailed description in conjunction with the accompanying drawings, wherein:

[0033] Figure 1 is a schematic conceptual diagram showing a neural processing unit according to the present disclosure;

[0034] Figure 2 is a schematic conceptual diagram showing one processing element of an array of processing elements suitable for the present disclosure;

[0035] Figure 3 shows Figure 1 an example diagram of a modified example of the neural processing unit 100;

[0036] Figure 4 is a schematic conceptual diagram showing an exemplary artificial neural network model;

[0037] Figure 5A is a view showing the basic structure of a convolutional neural network;

[0038] Figure 5B is a view showing the overall operation of a convolutional neural network.

[0039] Figure 6 shows including Figure 1 or an NPU of 3, a view of an exemplary architecture of a system-on-chip (SoC);

[0040] Figure 7 is a view showing an example of a scan flip-flop;

[0041] Figure 8 is a view showing an example of an architecture added for scan testing in a hardware design;

[0042] Figure 9A simply shows from an operational perspective Figure 6 an example diagram of the SoC;

[0043] Figure 9B is an example diagram showing a configuration for testing an NPU;

[0044] Figure 10 is a schematic diagram showing the operation of a wrapper;

[0045] Figure 11 is a schematic diagram showing the internal configuration of an ICT;

[0046] Figure 12 is a block diagram specifically showing the operation of monitoring whether an ICT functional component is in an idle state;

[0047] Figure 13It is an example diagram showing the operations among a host, a slave, and an arbiter operating on a system bus;

[0048] Figure 14 It is a view showing an example of adding a shift register in an SoC chip;

[0049] Figure 15 It is an example diagram showing the operation sequence of ICT;

[0050] Figure 16 It is a block diagram for facilitating the understanding of the test process of the shown internal memory;

[0051] Figure 17 It is an example diagram showing the process of testing a function using a random number generator;

[0052] Figure 18A It is a view showing an example of multiple clocks, Figure 18B It is an example diagram showing the operation of a tester under multiple clocks; Figure 18C It is a view showing the path of test input data;

[0053] Figure 19A It is a view showing an example of a functional component; Figure 19B It is a view showing an example of test input data (e.g., test vectors) imported into a tester in ICT;

[0054] Figure 20 It is a view showing the test process using DFT;

[0055] Figure 21 It is a view showing an example of shifted data and captured data during a test process;

[0056] Figure 22 It is a view showing an example of switching a test mode to a normal operation mode;

[0057] Figure 23 It is a view showing an example of a flip - flop operating on a scan chain;

[0058] Figure 24 It is a view showing a part of a CUT operating in a normal operation mode;

[0059] Figure 25 It is a view showing an example of a simulation process; and

[0060] Figure 26 It is a view showing the test architecture of a JPEG image encoder. Detailed Description of the Invention

[0061] The specific structures or step-by-step descriptions of the embodiments according to the concept of the present disclosure disclosed in this specification or application are only for the purpose of illustrating the embodiments according to the concept of the present disclosure. Examples according to the concept of the present disclosure can be implemented in various forms and are not to be construed as limited to the examples described in this specification or application.

[0062] Various modifications and variations can be applied to the examples according to the concept of the present disclosure, and the examples can have various forms. Thus, the examples will be described in detail in the specification or application with reference to the accompanying drawings. However, it should be understood that the examples according to the concept of the present disclosure are not limited to the specific examples, but include all variations, equivalents, or alternatives included within the spirit and technical scope of the present disclosure.

[0063] Terms such as first and / or second can be used to describe various components, but the components are not limited to the above terms. The above terms are used to distinguish one component from another. For example, without departing from the scope of the concept of the present invention, the first component can be referred to as the second component, and similarly, the second component can be referred to as the first component.

[0064] It should be understood that when describing that an element is "coupled" or "connected" to another element, the element can be directly coupled or directly connected to the other element, or can be coupled or connected to the other element through a third element. In contrast, when describing that an element is "directly coupled" or "directly connected" to another element, it should be understood that there is no element between them. Other expressions for describing the relationship between components, such as "interposed between", "adjacent", and "directly adjacent", should be interpreted in the same way.

[0065] The terms used in this specification are only for describing specific examples and are not intended to limit the present disclosure. If there is no obvious contrary meaning in the context, the singular form can include the plural form. In this specification, it should be understood that the term "comprises" or "has" means the presence of the features, quantities, steps, operations, components, parts, or combinations thereof described in the specification, but does not exclude the possibility of pre-existing or adding one or more other features, quantities, steps, operations, components, parts, or combinations thereof.

[0066] All terms, including technical or scientific terms, used herein have the same meanings as those commonly understood by those of ordinary skill in the art if there is no contrary definition. Terms defined in a general dictionary should be interpreted as having the same meaning as in the context of the relevant technology, but should not be interpreted as ideal or overly formal meanings if not clearly defined in this specification.

[0067] When describing examples, technologies that are well-known in the technical field of the present disclosure and have no direct relation to the present disclosure will not be described. The reason is that unnecessary descriptions are omitted to clearly convey the gist of the present disclosure without obscuring the gist.

[0068] <Term Definitions>

[0069] Here, to help understand the disclosure presented in this specification, terms used in this specification will be briefly defined.

[0070] NPU is an abbreviation for Neural Processing Unit, referring to a processor dedicated to operating an artificial neural network model separately from a Central Processing Unit (CPU).

[0071] ANN is an abbreviation for Artificial Neural Network, referring to a network that imitates human intelligence by connecting nodes in a hierarchical structure via synapses to simulate neurons in the human brain.

[0072] Information about the structure of an artificial neural network includes information about the number of layers, the number of nodes in each layer, information about the values of each node, information about the operation processing method, and information about the weight matrix applied to each node.

[0073] Information about the data locality of an artificial neural network is information for predicting the operation order of an artificial neural network model processed by a neural processing unit based on the data access request order in which the neural processing unit requests data from a separate memory.

[0074] DNN is an abbreviation for Deep Neural Network, which may mean increasing the number of hidden layers of an artificial neural network to achieve higher artificial intelligence.

[0075] CNN is an abbreviation for Convolutional Neural Network, which is a neural network whose function is similar to image processing performed in the visual cortex of the human brain. Convolutional neural networks are known to be applicable to image processing and are known to be easy to extract features of input data and recognize feature patterns.

[0076] A kernel refers to the weight matrix applied to a CNN.

[0077] Hereinafter, examples of the present disclosure will be described in detail by referring to the accompanying drawings to explain the present disclosure.

[0078] Figure 1 An illustration shows a neural processing unit according to the present disclosure.

[0079] Figure 1 The illustrated Neural Processing Unit (NPU) 100 is a processor dedicated to performing operations for an artificial neural network.

[0080] An artificial neural network refers to a network in which artificial neurons are aggregated. When there are various inputs or incoming stimuli, it multiplies the weights by the inputs or stimuli, sums the products, and then uses an activation function to convert the value after adding the bias for transmission. The artificial neural network trained as described above can be used to output an inference result based on input data.

[0081] The neural processing unit 100 can be a semiconductor device implemented by an electrical / electronic circuit. The electrical / electronic circuit can refer to a circuit including a large number of electronic components (transistors, capacitors, etc.). The neural processing unit 100 includes a processing element (PE) array 110, an NPU internal memory 120, an NPU scheduler 130, and an NPU interface 140. Each of the processing element array 110, the NPU internal memory 120, the NPU scheduler 130, and the NPU interface 140 can be a semiconductor circuit connected with a large number of electronic components. Therefore, some electronic components may be difficult to identify or distinguish with the naked eye and can only be identified through operation. For example, any circuit can operate as the processing element array 110 or as the NPU scheduler 130.

[0082] The neural processing unit 100 can include a processing element array 110, an NPU internal memory 120 configured to store an artificial neural network model inferred from the processing element array 110, and an NPU scheduler 130 configured to control the processing element array 110 and the NPU internal memory 120 based on data locality information of the artificial neural network model or information about its structure. Here, the artificial neural network model can include data locality information of the artificial neural network or information about its structure. The artificial neural network model can refer to an AI recognition model trained to perform a specific inference function.

[0083] The processing element array 110 can perform the operations of an artificial neural network. For example, when input data is input, the processing element array 110 can enable the artificial neural network to perform learning. When input data is input after learning is completed, the processing element array 110 can perform an operation of inferring an inference result through the artificial neural network that has completed learning.

[0084] The NPU interface 140 can communicate with various components (such as a memory) in the Figure 6 ANN driving device via a system bus.

[0085] For example, the neural processing unit 100 can call data of the artificial neural network model stored in the Figure 6 internal memory 400 to the NPU internal memory 120 through the NPU interface 140.

[0086] The NPU scheduler 130 is configured to control the operations of the processing element array 110 and the read / write instructions of the NPU internal memory 120 to perform inference operations of the neural processing unit 100.

[0087] The NPU scheduler 130 may be configured to analyze data locality information or information about its structure of an artificial neural network model to control the processing element array 110 and the NPU internal memory 120.

[0088] The NPU scheduler 130 may analyze or receive the structure of an artificial neural network model that can operate in the processing element array 110. The data of the artificial neural network that may be included in the artificial neural network model may store node data of each layer, layer placement data locality information or information about the structure, and weight data of each connection network connecting the nodes of the layers. The data of the artificial neural network may be stored in the memory provided in the NPU scheduler 130 or the NPU internal memory 120. The NPU scheduler 130 may access Figure 6 the internal memory 400 or the external memory to utilize the necessary data. However, not limited thereto, data locality information or information about its structure of the artificial neural network model can be generated based on data such as node data and weight data of the artificial neural network model. The weight data may also be referred to as a weight kernel. The node data is also called a feature map. For example, data defining the structure of the artificial neural network model may be generated when designing the artificial neural network model or completing learning, but the present disclosure is not limited thereto.

[0089] The NPU scheduler 130 may schedule the operation sequence of the artificial neural network model based on the data locality information or information about its structure of the artificial neural network model.

[0090] The NPU scheduler 130 may obtain the memory address values storing the node data of the layers of the artificial neural network model and the weight data of the connection network based on the data locality information or information about its structure of the artificial neural network model. For example, the NPU scheduler 130 may obtain the memory address values storing the node data of the layers of the artificial neural network model and the weight data of the connection network stored in the memory. Therefore, the NPU scheduler 130 may obtain the node data of the layers of the artificial neural network model to be driven and the weight data of the connection network from the internal memory 400 to store the obtained data in the NPU internal memory 120. The node data of each layer may have corresponding memory address values. The weight data of each connection network may have corresponding memory address values.

[0091] The NPU scheduler 130 may schedule the operation order of the processing element array 110 based on the data locality information of the artificial neural network model or information about its structure (such as the placement of the layers of the artificial neural network, the data locality information of the artificial neural network model, or information about its structure).

[0092] The NPU scheduler 130 schedules based on the data locality information of the artificial neural network model or information about its structure, such that the NPU scheduler can operate in a manner different from the scheduling concept of a general CPU. Considering fairness, efficiency, stability, and response time, the scheduling operation of a general CPU aims to provide the highest efficiency. That is, considering priorities and operation times, the general CPU scheduling executes the most processing within the same time.

[0093] Known CPUs use algorithms that consider data such as the priority of each process or the operation processing time to schedule tasks. In contrast, the NPU scheduler 130 can determine the processing order based on the data locality information of the artificial neural network model or information about its structure.

[0094] In addition, the NPU scheduler 130 can determine the processing order based on the data locality information of the artificial neural network model or information about its structure and / or the data locality information of the neural processing unit 100 to be used or information about its structure.

[0095] However, the present disclosure is not limited to the data locality information of the neural processing unit 100 or information about its structure. For example, the data locality information of the neural processing unit 100 or information about its structure can determine the processing order by using at least one of the memory size of the NPU internal memory 120, the hierarchical structure of the NPU internal memory 120, the number (size) data of the processing elements PE1 to PE12, and the operator structure of the processing elements PE1 to PE12. That is, the data locality information of the neural processing unit 100 or information about its structure can include at least one of the memory size of the NPU internal memory 120, the hierarchical structure of the NPU internal memory 120, the number data of the processing elements PE1 to PE12, and the operator structure of the processing elements PE1 to PE12. However, the present disclosure is not limited to the data locality information of the neural processing unit 100 or information about its structure. The memory size of the NPU internal memory 120 includes information about the memory capacity. The hierarchical structure of the NPU internal memory 120 includes information about the connection relationship between specific layers of each hierarchical structure. The operator structure of the processing elements PE1 to PE12 includes information about the components in the processing elements.

[0096] The neural processing unit 100 according to an example of the present disclosure may include at least one processing element, an NPU internal memory 120 that stores an artificial neural network model inferred from the at least one processing element, and an NPU scheduler 130 configured to control the at least one processing element and the NPU internal memory 120 based on data locality information of the artificial neural network model or information about its structure. The NPU scheduler 130 may be configured to be further provided with data locality information of the neural processing unit 100 or information about its structure. In addition, the data locality information of the neural processing unit 100 or information about its structure may include at least one of the memory size of the NPU internal memory 120, the hierarchical structure of the NPU internal memory 120, the number (size) data of at least one processing unit, and the operator structure of at least one processing unit.

[0097] According to the structure of the artificial neural network model, the operations of each layer are performed sequentially. That is, when the structure of the artificial neural network model is determined, the operation order of each layer can be determined. According to the structure of the artificial neural network model, the order of operations or the order of data flow can be defined as the data locality of the artificial neural network model at the algorithm level.

[0098] When the compiler compiles the artificial neural network model to be executed in the neural processing unit 100, the artificial neural network data locality at the neural processing unit-memory level of the artificial neural network model can be reconstructed.

[0099] That is, the data locality of the artificial neural network model at the neural processing unit-memory level can be constructed according to the compiler, the algorithm applied to the artificial neural network model, and the operation characteristics of the neural processing unit 100.

[0100] For example, even in the same artificial neural network model, the artificial neural network data locality of the artificial neural network model to be processed can be configured differently according to the method of operating the artificial neural network model through the neural processing unit 100, such as the feature map tiling or stationary technique of the processing element, the number of processing elements of the neural processing unit 100, the cache memory capacity in the neural processing unit 100 such as the feature map or weights, the memory hierarchical structure in the neural processing unit 100, and the algorithm characteristics of the compiler that determines the operation order of the neural processing unit 100 to operate the artificial neural network model. This is because even when calculating the same artificial neural network model through the above factors, the neural processing unit 100 can determine the order of required data differently at each moment in terms of clock cycles.

[0101] The compiler can determine the order of data required for physical operation processing by constructing the artificial neural network data locality of the artificial neural network model at the word level in the neural processing unit-memory level.

[0102] In other words, the artificial neural network data locality of the artificial neural network model existing at the neural processing unit-memory level can be defined as information for predicting the operation order of the artificial neural network model processed by the neural processing unit 100 based on the data access request order requested by the neural processing unit 100 from the internal memory 400.

[0103] The NPU scheduler 130 can be configured to store the data locality information of the artificial neural network or information about its structure.

[0104] That is, even by using only the data locality information of the artificial neural network of the artificial neural network model or information about its structure, the NPU scheduler 130 can determine the processing order (sequence). That is, the NPU scheduler 130 can determine the operation order by using the data locality information or information about the structure from the input layer to the output layer of the artificial neural network. For example, the input layer operation can be scheduled first, and the output layer operation can be scheduled last. Therefore, when the NPU scheduler 130 is provided with the data locality information of the artificial neural network model or information about its structure, all operation sequences of the artificial neural network model can be known. Therefore, all scheduling orders can be determined.

[0105] In addition, the NPU scheduler 130 can determine the processing order and optimize the processing for each determined order by considering the data locality information of the artificial neural network model or information about its structure and the data locality information of the neural processing unit 100 or information about its structure.

[0106] Therefore, when all the data locality information of the artificial neural network model or information about its structure and the data locality information of the neural processing unit 100 or information about its structure are provided to the NPU scheduler 130, the running efficiency of each scheduling order determined by the data locality information or information about the structure of the artificial neural network model can be further improved. For example, the NPU scheduler 130 can acquire connection network data having four artificial neural network layers and weight data of three layers connecting these layers. In this case, a method for the NPU scheduler 130 to schedule the processing order based on the data locality information of the artificial neural network model or information about its structure will be described by way of example below.

[0107] For example, the NPU scheduler 130 may set the input data for the inference operation to the node data of the first layer, which is the input layer of the artificial neural network model, and schedule to first perform the multiply-accumulate (MAC) operation on the node data of the first layer and the weight data of the first connection network corresponding to the first layer. However, the examples of the present disclosure are not limited to MAC operations and may perform artificial neural network operations by using multipliers and adders that can be modified in various forms. Hereinafter, for convenience of description, the corresponding operation is referred to as the first operation, the result of the first operation is referred to as the first operation value, and the corresponding schedule may be referred to as the first schedule.

[0108] For example, after the first schedule, the NPU scheduler 130 may set the first operation value to the node data of the second layer corresponding to the first connection network, and schedule to perform the MAC operation on the node data of the second layer and the weight data of the second connection network corresponding to the second layer. Hereinafter, for convenience of description, the corresponding operation is referred to as the second operation, the result of the second operation is referred to as the second operation value, and the corresponding schedule may be referred to as the second schedule.

[0109] For example, during the second schedule, the NPU scheduler 130 may set the second operation value to the node data of the third layer corresponding to the second connection network, and schedule to perform the MAC operation on the node data of the third layer and the weight data of the third connection network corresponding to the third layer. Hereinafter, for convenience of description, the corresponding operation is referred to as the third operation, the result of the third operation is referred to as the third operation value, and the corresponding schedule may be referred to as the third schedule.

[0110] For example, the NPU scheduler 130 may set the third operation value to the node data of the fourth layer, which is the output layer corresponding to the third connection network, and schedule to store the inference result stored in the node data of the fourth layer in the NPU internal memory 120. Hereinafter, for convenience of description, the corresponding schedule may be referred to as the fourth schedule.

[0111] In summary, the NPU scheduler 130 may control the NPU internal memory 120 and the processing element array 110 to operate in the order of the first schedule, the second schedule, the third schedule, and the fourth schedule. That is, the NPU scheduler 130 may be configured to control the NPU internal memory 120 and the processing element array 110 to perform operations according to the set schedule order.

[0112] In summary, the neural processing unit 100 according to the example of the present disclosure may be configured to schedule the processing order based on the structure of the layers of the artificial neural network and the operation order data corresponding to the structure.

[0113] For example, the NPU scheduler 130 may be configured to schedule the processing order from the input layer to the output layer of the artificial neural network of the artificial neural network model based on data locality information or information about the structure.

[0114] The NPU scheduler 130 can improve the operation speed and memory reusability of the neural processing unit by controlling the NPU internal memory 120 via the utilization of the scheduling order based on the data locality information of the artificial neural network model or the information about its structure.

[0115] According to the characteristics of the artificial neural network operation driven in the neural processing unit 100 according to an example of the present disclosure, the operation value of one layer can be used as the input data of the subsequent layer.

[0116] Therefore, the neural processing unit 100 controls the NPU internal memory 120 according to the scheduling order to improve the memory reusability of the NPU internal memory 120. The reuse of the memory can be determined by the number of times of reading the data stored in the memory. For example, after storing specific data in the memory, when the specific data is read only once and then deleted or overwritten, the memory reuse rate may become 100%. For example, after storing specific data in the memory, when the specific data is read four times and then deleted or overwritten, the memory reuse rate may become 400%. That is to say, the memory reuse rate can be determined as the number of times of multiplexing the stored data. That is to say, the memory reuse rate may refer to the reuse of the data stored in the memory or the specific memory address storing the specific data.

[0117] Specifically, when the NPU scheduler 130 is configured to be provided with the data locality information of the artificial neural network model or the information about its structure and calculate the sequential data, in this sequential data, the operation of the artificial neural network is performed through the provided data locality information of the artificial neural network model or the information about its structure, and the NPU scheduler 130 identifies the operation result of the node data of a specific layer of the artificial neural network model and the weight data of a specific connection network as the node data of the corresponding subsequent layer.

[0118] Therefore, the NPU scheduler 130 can reuse the value of the memory address storing the specific operation result for subsequent operations. Therefore, the memory reuse rate can be improved.

[0119] For example, set the first operation value of the first schedule as the node data of the second layer of the second schedule. More specifically, the NPU scheduler 130 may reset the memory address value corresponding to the first operation value of the first schedule stored in the NPU internal memory 120 to the memory address value corresponding to the node data of the second layer of the second schedule. That is to say, the memory address value can be reused. Therefore, the NPU scheduler 130 reuses the data of the memory address of the first schedule, so that the NPU internal memory 120 can utilize it as the node data of the second layer of the second schedule without a separate memory write operation.

[0120] For example, set the second operation value of the second schedule as the node data of the third layer of the third schedule. More specifically, the NPU scheduler 130 may reset the memory address value corresponding to the second operation value of the second schedule stored in the NPU internal memory 120 to the memory address value corresponding to the node data of the third layer of the third schedule. That is to say, the memory address value can be reused. Therefore, the NPU scheduler 130 reuses the data of the memory address of the second schedule, so that the NPU internal memory 120 can utilize the data as the node data of the third layer of the third schedule without a separate memory write operation.

[0121] For example, set the third operation value of the third schedule as the node data of the fourth layer of the fourth schedule. More specifically, the NPU scheduler 130 may reset the memory address value corresponding to the third operation value of the third schedule stored in the NPU internal memory 120 to the memory address value corresponding to the node data of the fourth layer of the fourth schedule. That is to say, the memory address value can be reused. Accordingly, the NPU scheduler 130 reuses the data of the memory address of the third schedule, so that the NPU internal memory 120 can utilize the data as the node data of the fourth layer of the fourth schedule without a separate memory write operation.

[0122] In addition, the NPU scheduler 130 may be configured to determine whether to reuse the scheduling order and memory to control the NPU internal memory 120. In this case, the NPU scheduler 130 analyzes the data locality information of the artificial neural network model or the information about its structure to provide an effective schedule. In addition, the data required for the operation that can reuse the memory is not repeatedly stored in the NPU internal memory 120, thereby reducing the memory usage. In addition, the NPU scheduler 130 can improve the efficiency of the NPU internal memory 120 by calculating the reduced memory usage as much as the memory is reused.

[0123] In addition, the NPU scheduler 130 can be configured to monitor the resource usage of the NPU internal memory 120 and the resource usage of the processing elements PE1 to PE12 based on the data locality information of the neural processing unit 100 or information about its structure. Therefore, the hardware resource utilization efficiency of the neural processing unit 100 can be improved.

[0124] The NPU scheduler 130 of the neural processing unit 100 according to an example of the present disclosure can have the effect of reusing memory by utilizing the data locality information of the artificial neural network model or information about its structure.

[0125] In other words, when the artificial neural network model is a deep neural network, the number of layers and the number of connected networks may increase significantly, and in this case, the effect of memory reuse can be maximized to a greater extent.

[0126] That is to say, when the neural processing unit 100 does not calculate the data locality information of the artificial neural network model or information about the structure and operation order, the NPU scheduler 130 may not determine whether to reuse the values stored in the NPU internal memory 120. Therefore, the NPU scheduler 130 unnecessarily generates the memory addresses required for each process, and copies substantially the same data from one memory address to another memory address. As a result, unnecessary memory read and write tasks are generated, and duplicate values are stored in the NPU internal memory 120, which may lead to the problem of unnecessary waste of memory.

[0127] The processing element array 110 refers to a configuration in which a plurality of processing elements PE1 to PE12 are provided, and the plurality of processing elements PE1 to PE12 are configured to operate on the node data of the artificial neural network and the weight data of the connection network. Each processing element may include a multiply-accumulate (MAC) operator and / or an arithmetic logic unit (ALU) operator, but is not limited thereto according to an example of the present disclosure.

[0128] Although Figure 2 A plurality of processing units are shown as an example, but operators implemented by a plurality of multipliers and adder trees can also be configured to be arranged in parallel in one processing unit instead of in the MAC. In this case, the processing element array 110 can also be referred to as at least one processing element including a plurality of operators.

[0129] The processing element array 110 is configured to include a plurality of processing elements PE1 to PE12. Figure 2The multiple processing elements PE1 to PE12 are merely examples for convenience of description, and the number of the multiple processing elements PE1 to PE12 is not limited. The size or number of the processing element array 110 can be determined by the number of the multiple processing elements PE1 to PE12. The size of the processing element array 110 can be implemented by an N×M matrix. Here, N and M are integers greater than zero. The processing element array 110 can include N×M processing elements. That is to say, one or more processing elements can be provided.

[0130] The size of the processing element array 110 can be designed in consideration of the characteristics of the artificial neural network model (in which the neural processing unit 100 operates). As an additional note, the number of processing elements can be determined in consideration of the data size on which the artificial neural network model runs, the required running speed, and the required power consumption. The data size of the artificial neural network model can be determined to correspond to the number of layers of the artificial neural network model and the weight data size of each layer.

[0131] Therefore, the size of the processing element array 110 of the neural processing unit 100 according to the examples of the present disclosure is not limited. As the number of processing elements of the processing element array 110 increases, the parallel computing ability of the operating artificial neural network model increases, but the manufacturing cost and physical size of the neural processing unit 100 may increase.

[0132] For example, the artificial neural network model running in the neural processing unit 100 can be an artificial neural network trained to detect thirty specific keywords, that is, an AI keyword recognition model. In this case, considering the characteristics of the operation amount, the size of the processing element array 110 of the neural processing unit 100 can be designed as 4×3. In other words, the neural processing unit 100 can include twelve processing elements. However, it is not limited thereto, and the number of the multiple processing elements PE1 to PE12 can be selected in the range of 8 to 16,384. That is to say, the examples of the present disclosure are not limited to the number of processing elements.

[0133] The processing element array 110 is configured to perform functions such as addition, multiplication, and accumulation required for artificial neural network operations. In other words, the processing element array 110 can be configured to perform multiply-accumulate (MAC) operations.

[0134] Hereinafter, the first processing element PE1 in the processing element array 110 will be taken as an example for description.

[0135] Figure 2 A processing element applicable to the processing element array of the present disclosure is illustrated.

[0136] The neural processing unit 100 according to an example of the present disclosure may include an array of processing elements 110, an NPU internal memory 120 configured to store an artificial neural network model inferred from the array of processing elements 110, and an NPU scheduler 130 configured to control the array of processing elements 110 and the NPU internal memory 120 based on data locality information of the artificial neural network model or information about its structure. The array of processing elements 110 is configured to perform MAC operations, and the array of processing elements 110 is configured to quantize and output the MAC operation results, but embodiments of the present disclosure are not limited thereto.

[0137] The NPU internal memory 120 may store all or part of the artificial neural network model according to the memory size and data size of the artificial neural network model.

[0138] The first processing element PE1 may include a multiplier 111, an adder 112, an accumulator 113, and a bit quantization unit 114. However, examples according to the present disclosure are not limited thereto, and the array of processing elements 110 may be modified in consideration of the operation characteristics of the artificial neural network.

[0139] The multiplier 111 multiplies the input (N)-bit data and (M)-bit data. The operation value of the multiplier 111 is output as (N+M)-bit data. Here, N and M are integers greater than zero. The first input unit that receives the (N)-bit data may be configured to receive a value with variable characteristics, and the second input unit that receives the (M)-bit data may be configured to receive a value with constant characteristics. When the NPU scheduler 130 differentiates between variable value and constant value characteristics, the NPU scheduler 130 may increase the memory reusability of the NPU internal memory 120. However, the input data of the multiplier 111 is not limited to constant values and variable values. That is, according to examples of the present disclosure, the input data of the processing element may be operated by understanding the characteristics of constant values and variable values, thereby improving the operation efficiency of the neural processing unit 100. However, the neural processing unit 100 is not limited to the characteristics of constant values and variable values of the input data.

[0140] Here, the value with the characteristic of a variable or the variable refers to a value of a memory address storing a corresponding value being updated whenever the input input data is updated. For example, the node data of each layer may be a MAC operation value reflecting the weight data of the artificial neural network model. When performing object recognition of moving image data or the like through the artificial neural network model, the input image changes at each frame, thereby changing the node data of each layer.

[0141] Here, regardless of the update of the input data, values with constant characteristics are stored, where the constant can refer to the value of the memory address storing the corresponding value. For example, even if object recognition of moving image data, etc., is inferred by an artificial neural network model based on the unique inference determination criteria of the artificial neural network model, the weight data of the connection network may not change.

[0142] That is to say, the multiplier 111 can be configured to receive a variable and a constant. As an additional explanation, the variable value input to the first input unit can be the node data of the layer of the artificial neural network, and the node data can be the input data of the input layer of the artificial neural network, the accumulated value of the hidden layer, and the accumulated value of the output layer. The constant value input to the second input unit can be the weight data of the connection network of the artificial neural network.

[0143] The NPU scheduler 130 can be configured to improve the memory reuse rate in consideration of the characteristics of the constant value.

[0144] The variable value is the operation value of each layer, and the NPU scheduler 130 can identify reusable variable values based on the data locality information of the artificial neural network model or information about its structure, and control the NPU internal memory 120 to reuse the memory.

[0145] The constant value is the weight data of each connection network, and the NPU scheduler 130 can identify the constant values of the connection networks that are reused based on the data locality information of the artificial neural network model or information about its structure, and control the NPU internal memory 120 to reuse the memory.

[0146] That is to say, the NPU scheduler 130 can be configured to identify reusable variable values and reusable constant values based on the data locality information of the artificial neural network model or information about its structure, and control the NPU internal memory 120 to reuse the memory.

[0147] The processing element knows that when 0 is input to one of the first input unit and the second input unit of the multiplier 111, even if no operation is performed, the operation result is 0. Therefore, the processing element can limit the operation of the multiplier 111 to not perform the operation.

[0148] For example, when 0 is input to one of the first input unit and the second input unit of the multiplier 111, the multiplier 111 can be configured to operate in a zero-jump method.

[0149] The bit widths of the data input to the first input unit and the second input unit can be determined according to the quantization of the node data and the weight data of each layer in the artificial neural network model. For example, when the node data of the first layer can be quantized to five bits and the weight data of the first layer can be quantized to seven bits. In this case, the first input unit can be configured to receive five-bit data, and the second input unit can be configured to receive seven-bit data.

[0150] When the quantized data stored in the NPU internal memory 120 is input to the input unit of the processing element, the neural processing unit 100 can control the quantized bit width to be converted in real time. That is, each layer can have a different quantized bit width, and when the bit width of the input data is converted, the processing element can be configured to receive the bit width information from the neural processing unit 100 in real time and convert the bit width in real time to generate the input data.

[0151] The accumulator 113 uses the adder 112 to accumulate the operation value of the multiplier 111 and the operation value of the accumulator 113 as many times as the number of (L) loops. Therefore, the bit widths of the data of the output unit and the input unit of the accumulator 113 can be output to (N+M+log2(L)) bits. Here, L is an integer greater than zero.

[0152] When the accumulation is completed, the accumulator 113 is applied with an initialization reset to initialize the data stored in the accumulator 113 to zero, but the example according to the present disclosure is not limited thereto.

[0153] The bit quantization unit 114 can reduce the bit width of the data output from the accumulator 113. The bit quantization unit 114 can be controlled by the NPU scheduler 130. The bit width of the quantized data can be output to (X) bits. Here, X is an integer greater than zero. According to the above configuration, the processing element array 110 is configured to perform MAC operations, and the processing element array 110 can quantize the MAC operation results to output the results. Quantization may have the effect that the larger the number of (L) loops, the lower the power consumption. In addition, when the power consumption is reduced, the heat generation can also be reduced. Specifically, when the heat generation is reduced, the possibility of the neural processing unit 100 malfunctioning due to high temperature can be reduced.

[0154] The output data (X) bits of the bit quantization unit 114 can be used as the node data of the subsequent layer or the input data of the convolution. When the artificial neural network model is quantized, the bit quantization unit 114 can be configured to be provided with the quantization information from the artificial neural network model. However, it is not limited thereto, and the NPU scheduler 130 can also be configured to extract the quantization information by analyzing the artificial neural network model. Therefore, the output data (X) bits are converted to the quantized bit width to be output to correspond to the quantized data size. The output data (X) bits of the bit quantization unit 114 can be stored in the NPU internal memory 120 with the quantized bit width.

[0155] The processing element array 110 of the neural processing unit 100 according to an example of the present disclosure includes a multiplier 111, an adder 112, an accumulator 113, and a bit quantization unit 114. The processing element array 110 can reduce the data with a bit width of (N+M+log2(L)) bits output from the accumulator 113 to a bit width of (X) bits through the bit quantization unit 114. The NPU scheduler 130 can control the bit quantization unit 114 to reduce the bit width of the output data by a predetermined number of bits from the least significant bit (LSB) to the most significant bit (MSB). When the bit width of the output data decreases, power consumption, the amount of operations, and memory usage can be reduced. However, when the bit width is reduced below a specific length, there may be a problem that the inference accuracy of the artificial neural network model deteriorates sharply. Therefore, the reduction of the bit width of the output data, that is, the quantization level, can be determined by comparing the reduction amounts of power consumption, the amount of operations, and memory usage with the reduction level of the inference accuracy of the artificial neural network model. The quantization level can be determined by determining the target inference accuracy of the artificial neural network model and testing by gradually reducing the bit width. The quantization level can be determined for each operation value of each layer.

[0156] According to the above first processing element PE1, the processing element array 110 can increase the MAC operation speed by adjusting the bit widths of the (N) - bit data and (M) - bit data of the multiplier 111 and reduce the bit width of the operation value (X) bits through the bit quantization unit 114 while reducing power consumption. In addition, the convolution operation of the artificial neural network can be performed more efficiently.

[0157] The NPU internal memory 120 of the neural processing unit 100 can be a memory system configured in consideration of the MAC operation characteristics and power consumption characteristics of the processing element array 110.

[0158] For example, the neural processing unit 100 can be configured to reduce the bit width of the operation value of the processing element array 110 in consideration of the MAC operation characteristics and power consumption characteristics of the processing element array 110.

[0159] The NPU internal memory 120 of the neural processing unit 100 can be configured to minimize the power consumption of the neural processing unit 100.

[0160] The NPU internal memory 120 of the neural processing unit 100 can be a memory system configured to control the memory with low power in consideration of the data size and operation steps of the ongoing artificial neural network model.

[0161] The NPU internal memory 120 of the neural processing unit 100 can be a low - power memory system, which is configured to reuse a specific memory address storing weight data in consideration of the data size and operation steps of the ongoing artificial neural network model.

[0162] The neural processing unit 100 can provide various activation functions to impart non-linearity. For example, the neural processing unit 100 can provide a sigmoid function, a hyperbolic tangent function, or a ReLU function. The activation function can be selectively applied after the MAC operation. The operation value to which the activation function is applied can be referred to as an activation map.

[0163] Figure 3 is illustrated Figure 1 a modified example of the neural processing unit 100.

[0164] Figure 3 The neural processing unit 100 of Figure 1 is substantially the same as the processing unit 100 exemplarily shown in

[0165] Figure 1 except that the processing element array 110. For the sake of convenience of description, redundant descriptions will be omitted.

[0166] Figure 3 The multiple processing elements PE1 to PE12 and the multiple register files RF1 to RF12 of

[0167] are merely examples for convenience of description, and the number of the multiple processing elements PE1 to PE12 and the number of the multiple register files RE1 to RE12 are not limited.

[0168] The size or number of the processing element array 110 can be determined by the number of the multiple processing elements PE1 to PE12 and the number of the multiple register files RF1 to RF12. The size of the processing element array 110 and the multiple register files RF1 to RF12 can be implemented by an N×M matrix. Here, N and M are integers greater than zero.

[0169] The register files RF1 to RF12 of the neural processing unit 100 are static memory units directly connected to the processing elements PE1 to PE12. For example, the register files RF1 to RF12 may be configured by flip-flops and / or latches. The register files RF1 to RF12 may be configured to store the MAC operation values of the corresponding processing elements RF1 to RF12. The register files RF1 to RF12 may be configured to provide weight data and / or node data to the NPU system memory 120 or to be provided with weight data and / or node data from the NPU system memory 120.

[0170] Figure 4 An exemplary artificial neural network (ANN) model is illustrated.

[0171] Hereinafter, the operation of an exemplary artificial neural network model 110a that can be operated in the neural processing unit 100 will be explained.

[0172] Figure 4 The exemplary artificial neural network model 110a may be an artificial neural network trained in the neural processing unit 100 or in a separate machine learning device. The artificial neural network model 110a may be an artificial neural network trained to perform various inference functions such as object recognition or speech recognition.

[0173] The artificial neural network model 110a may be a deep neural network (DNN).

[0174] However, the artificial neural network model 110a according to an example of the present disclosure is not limited to a deep neural network.

[0175] For example, the artificial neural network model 110a may be a model such as a fully convolutional network (FCN) having VGG, VGG16, DenseNET, and encoder-decoder structures, a deep neural network (DNN), such as SegNet, DeconvNet, DeepLAB, V3+, or U-net, or SqueezeNet, Alexnet, ResNet18, MobileNet-v2, GoogLeNet, Resnet-v2, Resnet50, Resnet101, and Inception-v3, but the present disclosure is not limited thereto. In addition, the artificial neural network model 110a may be an ensemble model based on at least two different models.

[0176] The artificial neural network model 110a may be stored in the NPU internal memory 120 of the neural processing unit 100. Alternatively, the artificial neural network model 110a may be implemented as stored in Figure 6 the device 1000, or Figure 6in the internal memory 400 of the device 1000 and then loaded into the neural processing unit 100 during the operation of the artificial neural network model 110a.

[0177] Hereinafter, reference will be made to Figure 4 describe the inference process of the exemplary artificial neural network model 110a executed by the neural processing unit 100.

[0178] The artificial neural network model 110a can be an exemplary deep neural network model, which includes an input layer 110a-1, a first connection network 110a-2, a first hidden layer 110a-3, a second connection network 110a-4, a second hidden layer 110a-5, a third connection network 110a-6, and an output layer 110a-7. However, the present disclosure is not limited to only Figure 4 the artificial neural network model shown in. The first hidden layer 110a-3 and the second hidden layer 110a-5 can also be referred to as multiple hidden layers.

[0179] The input layer 110a-1 can exemplarily include input nodes x1 and x2. That is, the input layer 110a-1 can include information about two input values. Figure 1 The NPU scheduler 130 shown in FIG. 2 or 3 can set a memory address in Figure 1 the NPU internal memory 120 of FIG. 2 or 3, and store information about the input values from the input layer 110a-1 at this memory address.

[0180] For example, the first connection network 110a-2 can include information about six weight values for connecting the nodes of the input layer 110a-1 to the nodes of the first hidden layer 110a-3 respectively. Figure 1 The NPU scheduler 130 shown in FIG. 2 or 3 can set a memory address in the NPU internal memory 120, and store information about the weight values of the first connection network 110a-2 at this memory address. Each weight value is multiplied by the input node value, and the accumulated value of the multiplied values is stored in the first hidden layer 110a-3.

[0181] For example, the first hidden layer 110a-3 can include nodes a1, a2, and a3. That is, the first hidden layer 110a-3 can include information about three node values. Figure 1 The NPU scheduler 130 shown in FIG. 2 or 3 sets a memory address in the NPU internal memory 120 for storing information about the node values of the first hidden layer 110a-3.

[0182] For example, the second connection network 110a-4 can include information about nine weight values for connecting the nodes of the first hidden layer 110a-3 to the nodes of the second hidden layer 110a-5 respectively. Figure 1The NPU scheduler 130 of 3 or can set a memory address for storing information about the weight values of the second connection network 110a-4 in the NPU internal memory 120. The weight values of the second connection network 110a-4 are multiplied by the node values input from the corresponding first hidden layer 110a-3, and the accumulated value of the multiplied values is stored in the second hidden layer 110a-5.

[0183] For example, the second hidden layer 110a-5 may include nodes b1, b2, and b3. That is, the second hidden layer 110a-5 may include information about three node values. The NPU scheduler 130 can set a memory address in the NPU internal memory 120 for storing information about the node values of the second hidden layer 110a-5.

[0184] For example, the third connection network 110a-6 may include information about six weight values respectively connecting the nodes of the second hidden layer 110a-5 and the nodes of the output layer 110a-7. The NPU scheduler 130 can set a memory address for storing information about the weight values of the third connection network 110a-6 in the NPU internal memory 120. The weight values of the third connection network 110a-6 are multiplied by the node values input from the second hidden layer 110a-5, and the accumulated value of the multiplied values is stored in the output layer 110a-7.

[0185] For example, the output layer 110a-7 may include nodes y1 and y2. That is, the output layer 110a-7 may include information about two node values. The NPU scheduler 130 can set a memory address for storing information about the node values of the output layer 110a-7 in the NPU internal memory 120.

[0186] That is, the NPU scheduler 130 can analyze or receive the structure of an artificial neural network model that can operate in the processing element array 110. The information of the artificial neural network that can be included in the artificial neural network model may include: information about the node values of each layer, layer placement data locality information or information about the structure, and information about the weight values of each connection network connecting the nodes of the connection layers.

[0187] The NPU scheduler 130 is provided with data locality information or information about its structure of the exemplary artificial neural network model 110a, so that the NPU scheduler 130 can determine the operation order from the input to the output of the artificial neural network model 110a.

[0188] Therefore, considering the scheduling order, the NPU scheduler 130 can set the memory addresses in the NPU internal memory 120 that store the MAC operation values for each layer. For example, the specific memory addresses can be the MAC operation values of the input layer 110a-1 and the first connection network 110a-2, or the input data of the first hidden layer 110a-3. However, the present disclosure is not limited to MAC operation values, and the MAC operation values can also be referred to as artificial neural network operation values.

[0189] At this time, the NPU scheduler 130 knows that the MAC operation results of the input layer 110a-1 and the first connection network 110a-2 are the inputs to the first hidden layer 110a-3, so it is controlled to use the same memory address. That is, the NPU scheduler 130 can reuse the MAC operation values based on the data locality information of the artificial neural network model or the information about its structure. Therefore, the memory reuse function of the NPU system memory 120 can be provided.

[0190] That is, the NPU scheduler 130 can store the MAC operation values of the artificial neural network model 110a in a specific area specified by any memory address in the NPU internal memory 120 according to the scheduling order, and use the MAC operation values as the input data for the MAC operation in the subsequent scheduling order in the specific area where the MAC operation values are stored.

[0191] Viewing the MAC operation from the perspective of the first processing element PE1

[0192] The MAC operation will be described in detail from the perspective of the first processing element PE1. The first processing element PE1 can be specified to perform the MAC operation of the node a1 in the first hidden layer 110a-3.

[0193] First, the first processing element PE1 inputs the data of the node x1 in the input layer 110a-1 into the first input unit of the multiplier 111, and inputs the weight data between the node x1 and the node a1 into the second input unit. The adder 112 adds the operation value of the multiplier 111 and the operation value of the accumulator 113. At this time, when the count value of (L) loops is 0, there is no accumulated value, so the accumulated value is 0. Therefore, the count value of the operation value adder 112 can be equal to the operation value of the multiplier 111. At this time, the count value of (L) loops can be 1.

[0194] Next, the first processing element PE1 inputs the node x2 data of the input layer 110a-1 to the first input unit of the multiplier 111, and inputs the weight data between the node x2 and the node a1 to the second input unit. The adder 112 adds the operation value of the multiplier 111 and the operation value of the accumulator 113. At this time, when the value of (L) loops is 1, the node x1 data calculated in the previous step and the multiplication value of the weight between the node x1 and the node a1 are stored. Therefore, the adder 112 generates the MAC operation value of the node x1 and the node x2 corresponding to the node a1.

[0195] Third, the NPU scheduler 130 can complete the MAC operation of the first processing unit PE1 based on the data locality information of the artificial neural network model or the information about its structure. At this time, an initialization reset is input to initialize the accumulator 113. That is, the counter value of (L) loops can be initialized to 0.

[0196] The bit quantization unit 114 can be appropriately adjusted according to the accumulated value. In other words, as the value of (L) loops increases, the bit width of the output value increases. At this time, the NPU scheduler 130 can remove a predetermined lower bit so that the bit width of the operation value of the first processing element PE1 is (x) bits.

[0197] Viewing the MAC operation from the perspective of the second processing element PE2

[0198] The MAC operation will be described in detail from the perspective of the second processing element PE2. The second processing element PE2 can be specified to perform the MAC operation of the node a2 of the first hidden layer 110a-3.

[0199] First, the second processing element PE2 inputs the node xl data of the input layer 110a-1 to the first input unit of the multiplier 111, and inputs the weight data between the node xl and the node a2 to the second input unit. The adder 112 adds the operation value of the multiplier 111 and the operation value of the accumulator 113. At this time, when the value of (L) loops is 0, there is no accumulated value, so the accumulated value is 0. Therefore, the operation value of the adder 112 can be equal to the operation value of the multiplier 111. At this time, the counter value of (L) loops can be 1.

[0200] Second, the second processing element PE2 inputs the node x2 data of the input layer 110a-1 to the first input unit of the multiplier 111, and inputs the weight data between the node x2 and the node a2 to the second input unit. The adder 112 adds the operation value of the multiplier 111 and the operation value of the accumulator 113. At this time, when the value of (L) loops is 1, the node x1 data calculated in the previous step and the multiplication value of the weight between the node x1 and the node a2 are stored. Therefore, the adder 112 generates the MAC operation value of the node x1 and the node x2 corresponding to the node a2.

[0201] Third, the NPU scheduler 130 can complete the MAC operation of the first processing element PE1 based on the data locality information of the artificial neural network model or the information about its structure. At this time, an initialization reset is input to initialize the accumulator 113. That is, the counter value of (L) loops can be initialized to 0. The bit quantization unit 114 can be appropriately adjusted according to the accumulated value.

[0202] Viewing the MAC operation from the perspective of the third processing element PE3

[0203] The MAC operation will be described in detail from the perspective of the third processing element PE3. It can be specified that the third processing element PE3 performs the MAC operation of the node a3 of the first hidden layer 110a-3.

[0204] First, the third processing element PE3 inputs the node xl data of the input layer 110a-1 to the first input unit of the multiplier 111, and inputs the weight data between the node xl and the node a3 to the second input unit. The adder 112 adds the operation value of the multiplier 111 and the operation value of the accumulator 113. At this time, when the value of (L) loops is 0, there is no accumulated value, so the accumulated value is 0. Therefore, the operation value of the adder 112 can be equal to the operation value of the multiplier 111. At this time, the counter value of (L) loops can be 1.

[0205] Second, the third processing element PE3 inputs the node x2 data of the input layer 110a-1 to the first input unit of the multiplier 111, and inputs the weight data between the node x2 and the node a3 to the second input unit. The adder 112 adds the operation value of the multiplier 111 and the operation value of the accumulator 113. At this time, when the value of (L) loops is 1, the node x1 data calculated in the previous step and the multiplication value of the weight between the node x1 and the node a3 are stored. Therefore, the adder 112 generates the MAC operation value of the node x1 and the node x2 corresponding to the node a3.

[0206] Third, the NPU scheduler 130 may complete the MAC operation of the first processing element PE1 based on the data locality information of the artificial neural network model or the information about its structure. At this time, an initialization reset is input to initialize the accumulator 113. That is, the counter value of (L) loops may be initialized to 0. The bit quantization unit 114 may be appropriately adjusted according to the accumulated value.

[0207] Therefore, the NPU scheduler 130 of the neural processing unit 100 may execute the MAC operation of the first hidden layer 110a-3 by simultaneously using three processing elements PE1 to PE3.

[0208] Viewing the MAC operation from the perspective of the fourth processing element PE4

[0209] The MAC operation will be described in detail from the perspective of the fourth processing element PE4. The fourth processing element PE4 may be specified to execute the MAC operation of the node b1 of the second hidden layer 110a-5.

[0210] First, the fourth processing element PE4 inputs the node a1 data of the first hidden layer 110a-3 to the first input unit of the multiplier 111 and inputs the weight data between the node a1 and the node b1 to the second input unit. The adder 112 adds the operation value of the multiplier 111 and the operation value of the accumulator 113. At this time, when the (L) loops are 0, there is no accumulated value, so the accumulated value is 0. Therefore, the operation value of the adder 112 may be equal to the operation value of the multiplier 111. At this time, the counter value of the (L) loops may be 1.

[0211] Second, the fourth processing element PE4 inputs the node a2 data of the first hidden layer 110a-3 to the first input unit of the multiplier 111 and inputs the weight data between the node a2 and the node b1 to the second input unit. The adder 112 adds the operation value of the multiplier 111 and the operation value of the accumulator 113. At this time, when the (L) loops are 1, the node a1 data calculated in the previous step and the multiplied value of the weight between the node a1 and the node b1 are stored. Therefore, the adder 112 generates the MAC operation value of the node a1 and the node a2 corresponding to the node b1. At this time, the counter value of the (L) loops may be 2.

[0212] Third, the fourth processing element PE4 inputs the node a3 data of the input layer 110a-1 to the first input unit of the multiplier 111, and inputs the weight data between node a3 and node b1 to the second input unit. The adder 112 adds the operation value of the multiplier 111 and the operation value of the accumulator 113. At this time, when the number of (L) cycles is 2, the MAC operation values of node a1 and node a2 corresponding to node b1 calculated in the previous step are stored. Therefore, the adder 112 generates the MAC operation values of node a1, node a2, and node a3 corresponding to node b1.

[0213] Fourth, the NPU scheduler 130 can complete the MAC operation of the first processing element PE1 based on the data locality information of the artificial neural network model or the information about its structure. At this time, an initialization reset is input to initialize the accumulator 113. That is, the counter value of the (L) cycles can be initialized to 0. The bit quantization unit 114 can be appropriately adjusted according to the accumulated value.

[0214] Viewing the MAC operation from the perspective of the fifth processing element PE5

[0215] The MAC operation will be described in detail from the perspective of the fifth processing element PE5. The fifth processing element PE5 can be specified to perform the MAC operation of node b2 in the second hidden layer 110a-5.

[0216] First, the fifth processing element PE5 inputs the node a1 data of the first hidden layer 110a-3 to the first input unit of the multiplier 111, and inputs the weight data between node a1 and node b2 to the second input unit. The adder 112 adds the operation value of the multiplier 111 and the operation value of the accumulator 113. At this time, when the number of (L) cycles is 0, there is no accumulated value, so the accumulated value is 0. Therefore, the operation value of the adder 112 can be equal to the operation value of the multiplier 111. At this time, the counter value of the (L) cycles can be 1.

[0217] Second, the fifth processing element PE5 inputs the node a2 data of the first hidden layer 110a-3 to the first input unit of the multiplier 111, and inputs the weight data between node a2 and node b2 to the second input unit. The adder 112 adds the operation value of the multiplier 111 and the operation value of the accumulator 113. At this time, when the number of (L) cycles is 1, the node a1 data calculated in the previous step and the multiplication value of the weight between node a1 and node b2 are stored. Therefore, the adder 112 generates the MAC operation value of node a1 and node a2 corresponding to node b2. At this time, the counter value of the (L) cycles can be 2.

[0218] Third, the fifth processing element PE5 inputs the node a3 data of the input layer 110a-1 to the first input unit of the multiplier 111, and inputs the weight data between node a3 and node b2 to the second input unit. The adder 112 adds the operation value of the multiplier 111 and the operation value of the accumulator 113. At this time, when the number of (L) loops is 2, the MAC operation values of node a1 and node a2 corresponding to node b2 calculated in the previous step are stored. Therefore, the adder 112 generates the MAC operation values of node a1, node a2, and node a3 corresponding to node b2.

[0219] Fourth, the NPU scheduler 130 can complete the MAC operation of the first processing element PE1 based on the data locality information of the artificial neural network model or the information about its structure. At this time, an initialization reset is input to initialize the accumulator 113. That is, the counter value of the (L) loops can be initialized to 0. The bit quantization unit 114 can be appropriately adjusted according to the accumulated value.

[0220] Viewing the MAC operation from the perspective of the sixth processing element PE6

[0221] The MAC operation will be described in detail from the perspective of the sixth processing element PE6. The sixth processing element PE6 can be specified to perform the MAC operation of node b3 in the second hidden layer 110a-5.

[0222] First, the sixth processing element PE6 inputs the node a1 data of the first hidden layer 110a-3 to the first input unit of the multiplier 111, and inputs the weight data between node a1 and node b3 to the second input unit. The adder 112 adds the operation value of the multiplier 111 and the operation value of the accumulator 113. At this time, when the number of (L) loops is 0, there is no accumulated value, so the accumulated value is 0. Therefore, the operation value of the adder 112 can be equal to the operation value of the multiplier 111. At this time, the counter value of the (L) loops can be 1.

[0223] Second, the sixth processing element PE6 inputs the node a2 data of the first hidden layer 110a-3 to the first input unit of the multiplier 111, and inputs the weight data between node a2 and node b3 to the second input unit. The adder 112 adds the operation value of the multiplier 111 and the operation value of the accumulator 113. At this time, when the number of (L) loops is 1, the node a1 data calculated in the previous step and the multiplication value of the weight between node a1 and node b3 are stored. Therefore, the adder 112 generates the MAC operation value of node a1 and node a2 corresponding to node b3. At this time, the counter value of the (L) loops can be 2.

[0224] Third, the sixth processing element PE6 inputs the node a3 data of the input layer 110a-1 to the first input unit of the multiplier 111, and inputs the weight data between the node a3 and the node b3 to the second input unit. The adder 112 adds the operation value of the multiplier 111 and the operation value of the accumulator 113. At this time, when the number of (L) cycles is 2, the MAC operation values of the node a1 and the node a2 corresponding to the node b3 calculated in the previous step are stored. Therefore, the adder 112 generates the MAC operation values of the node a1, the node a2, and the node a3 corresponding to the node b3.

[0225] Fourth, the NPU scheduler 130 can complete the MAC operation of the first processing element PE1 based on the data locality information of the artificial neural network model or the information about its structure. At this time, an initialization reset is input to initialize the accumulator 113. That is, the counter value of the (L) cycles can be initialized to 0. The bit quantization unit 114 can be appropriately adjusted according to the accumulated value.

[0226] Therefore, the NPU scheduler 130 of the neural processing unit 100 can perform the MAC operation of the second hidden layer 110a-5 by simultaneously using three processing elements PE4 to PE6.

[0227] Viewing the MAC operation from the perspective of the seventh processing element PE7

[0228] The MAC operation will be described in detail from the perspective of the seventh processing element PE7. The seventh processing element PE7 can be specified to perform the MAC operation of the node y1 of the output layer 110a-7.

[0229] First, the seventh processing element PE7 inputs the node b1 data of the second hidden layer 110a-5 to the first input unit of the multiplier 111, and inputs the weight data between the node b1 and the node y1 to the second input unit. The adder 112 adds the operation value of the multiplier 111 and the operation value of the accumulator 113. At this time, when the number of (L) cycles is 0, there is no accumulated value, so the accumulated value is 0. Therefore, the operation value of the adder 112 can be equal to the operation value of the multiplier 111. At this time, the counter value of the (L) cycles can be 1.

[0230] Second, the seventh processing element PE7 inputs the node b2 data of the second hidden layer 110a-5 to the first input unit of the multiplier 111, and inputs the weight data between the node b2 and the node y1 to the second input unit. The adder 112 adds the operation value of the multiplier 111 and the operation value of the accumulator 113. At this time, when the (L) loop is 1, the node b1 data calculated in the previous step and the multiplication value of the weight between the node b1 and the node y1 are stored. Therefore, the adder 112 generates the MAC operation value of the node b1 and the node b2 corresponding to the node y1. At this time, the counter value of the (L) loop can be 2.

[0231] Third, the seventh processing element PE7 inputs the node b3 data of the input layer 110a-1 to the first input unit of the multiplier 111, and inputs the weight data between the node b3 and the node y1 to the second input unit. The adder 112 adds the operation value of the multiplier 111 and the operation value of the accumulator 113. At this time, when the (L) loop is 2, the MAC operation value of the node b1 and the node b2 corresponding to the node y1 calculated in the previous step is stored. Therefore, the adder 112 generates the MAC operation value of the node b1, the node b2, and the node b3 corresponding to the node y1.

[0232] Fourth, the NPU scheduler 130 can complete the MAC operation of the first processing element PE1 based on the data locality information of the artificial neural network model or the information about its structure. At this time, an initialization reset is input to initialize the accumulator 113. That is, the counter value of the (L) loop can be initialized to 0. The bit quantization unit 114 can be appropriately adjusted according to the accumulated value.

[0233] Viewing the MAC operation from the perspective of the eighth processing element PE8

[0234] The MAC operation will be described in detail from the perspective of the eighth processing element PE8. The eighth processing element PE8 can be specified to perform the MAC operation of the node y2 of the output layer 110a-7.

[0235] First, the eighth processing element PE8 inputs the node b1 data of the second hidden layer 110a-5 to the first input unit of the multiplier 111, and inputs the weight data between the node b1 and the node y2 to the second input unit. The adder 112 adds the operation value of the multiplier 111 and the operation value of the accumulator 113. At this time, when the (L) loop is 0, there is no accumulated value, so the accumulated value is 0. Therefore, the operation value of the adder 112 can be equal to the operation value of the multiplier 111. At this time, the counter value of the (L) loop can be 1.

[0236] Second, the eighth processing element PE8 inputs the node b2 data of the second hidden layer 110a-5 to the first input unit of the multiplier 111, and inputs the weight data between the node b2 and the node y2 to the second input unit. The adder 112 adds the operation value of the multiplier 111 and the operation value of the accumulator 113. At this time, when the value of (L) loops is 1, the node b1 data calculated in the previous step and the multiplication value of the weight between the node b1 and the node y2 are stored. Therefore, the adder 112 generates the MAC operation value of the node b1 and the node b2 corresponding to the node y2. At this time, the counter value of (L) loops can be 2.

[0237] Third, the eighth processing element PE8 inputs the node b3 data of the input layer 110a-1 to the first input unit of the multiplier 111, and inputs the weight data between the node b3 and the node y2 to the second input unit. The adder 112 adds the operation value of the multiplier 111 and the operation value of the accumulator 113. At this time, when the value of (L) loops is 2, the MAC operation value of the node b1 and the node b2 corresponding to the node y2 calculated in the previous step is stored. Therefore, the adder 112 generates the MAC operation value of the node b1, the node b2, and the node b3 corresponding to the node y2.

[0238] Fourth, the NPU scheduler 130 can complete the MAC operation of the first processing element PE1 based on the data locality information of the artificial neural network model or the information about its structure. At this time, an initialization reset is input to initialize the accumulator 113. That is, the counter value of (L) loops can be initialized to 0. The bit quantization unit 114 can be appropriately adjusted according to the accumulated value.

[0239] Therefore, the NPU scheduler 130 of the neural processing unit 100 can perform the MAC operation of the output layer 110a-7 by simultaneously using two processing elements PE7 and PE8.

[0240] When the MAC operation of the eighth processing element PE8 is completed, the inference operation of the artificial neural network model 110a can be completed. That is, the artificial neural network model 110a can determine that the inference operation of one frame is completed. If the neural processing unit 100 performs real-time inference on moving image data, the image data of the subsequent frame can be input to the input nodes x1 and x2 of the input layer 110a-1. At this time, the NPU scheduler 130 can store the image data of the subsequent frame in the memory address storing the input data of the input layer 110a-1. When this process is repeated for each frame, the neural processing unit 100 can perform the inference operation in real time. In addition, the already set memory address can be reused.

[0241] According to Figure 4Regarding the summary of the artificial neural network model 110a, the NPU scheduler 130 of the neural processing unit 100 can determine the operation scheduling order based on the data locality information of the artificial neural network model 110a or information about its structure for the inference operations of the artificial neural network model 110a. The NPU scheduler 130 can set the memory addresses required by the NPU internal memory 120 based on the operation scheduling order. The NPU scheduler 130 can set the memory addresses of the reuse memory based on the data locality information of the artificial neural network model 110a or information about its structure. The NPU scheduler 130 designates the processing elements PE1 to PE8 required for the inference operations to perform the inference operations.

[0242] In other words, when the number of weight data connected to a node increases to L, the number of (L) cycles of the accumulator of the processing element can be set to L - 1. That is, even if the weight data of the artificial neural network increases, the accumulator increases the number of times the accumulator is accumulated to easily perform the inference operations.

[0243] That is, the NPU scheduler 130 of the neural processing unit 100 according to an example of the present disclosure can control the processing element array 110 and the NPU internal memory 120 based on the data locality information of the artificial neural network model and information about its structure (including the data locality information and information about its structure of the input layer 110a-1, the first connection network 110a-2, the first hidden layer 110a-3, the second connection layer 110a-4, the second hidden layer 110a-5, the third connection layer 110a-6, and the output layer 110a-7).

[0244] That is, the NPU scheduler 130 can set the memory address values corresponding to the node data of the input layer 110a-1, the weight data of the first connection network 110a-2, the node data of the first hidden layer 110a-3, the weight data of the second connection layer 110a-4, the node data of the second hidden layer 110a-5, the weight data of the third connection layer 110a-6, and the node data of the output layer 110a-7 in the NPU storage system 110

[0245] Hereinafter, the scheduling of the NPU scheduler 130 will be described in detail. The NPU scheduler 130 can schedule the operation order of the artificial neural network model based on the data locality information of the artificial neural network model or information about its structure.

[0246] The NPU scheduler 130 can obtain the memory address values storing the node data of the layers of the artificial neural network model and the weight data of the connection network based on the data locality information of the artificial neural network model or information about its structure.

[0247] For example, the NPU scheduler 130 may obtain the memory address values of the node data of the layers of the artificial neural network model stored in the main memory and the weight data of the connection network. Thus, the NPU scheduler 130 may obtain from the main memory the node data of the layer of the artificial neural network model to be driven and the weight data of the connection network to store the data in the NPU internal memory 120. The node data of each layer may have a corresponding memory address value. The weight data of each connection network may have a corresponding memory address value.

[0248] The NPU scheduler 130 may schedule the operation order of the processing element array 110 based on the data locality information of the artificial neural network model or information about its structure (e.g., the placement data locality information of the artificial neural network layers of the artificial neural network model or information about its structure).

[0249] For example, the NPU scheduler 130 may obtain weight data having four artificial neural network layers and the weight values of three layers connecting these layers, i.e., connection network data. In this case, a method of scheduling the processing order by the NPU scheduler 130 based on the data locality information of the artificial neural network model or information about its structure will be described by way of example below.

[0250] For example, the NPU scheduler 130 may set the input data for the inference operation as the node data of the first layer of the input layer 110a-1 of the artificial neural network model 110a and schedule to first perform the multiply-accumulate (MAC) operation on the node data of the first layer and the weight data of the first connection network corresponding to the first layer. Hereinafter, for convenience of description, the corresponding operation is referred to as the first operation, the result of the first operation is referred to as the first operation value, and the corresponding scheduling is referred to as the first scheduling.

[0251] For example, the NPU scheduler 130 may set the first operation value as the node data of the second layer corresponding to the first connection network after the first scheduling and schedule to perform the MAC operation on the node data of the second layer and the weight data of the second connection network corresponding to the second layer. Hereinafter, for convenience of description, the corresponding operation is referred to as the second operation, the result of the second operation is referred to as the second operation value, and the corresponding scheduling may be referred to as the second scheduling.

[0252] For example, the NPU scheduler 130 may set the second operation value as the node data of the third layer corresponding to the second connection network during the second scheduling and schedule to perform the MAC operation on the node data of the third layer and the weight data of the third connection network corresponding to the third layer. Hereinafter, for convenience of description, the corresponding operation is referred to as the third operation, the result of the third operation is referred to as the third operation value, and the corresponding scheduling is referred to as the third scheduling.

[0253] For example, the NPU scheduler 130 may set the third operation value as the node data of the fourth layer (the fourth layer is the output layer 110a-7 corresponding to the third connection network), and schedule to store the inference result in the node data of the fourth data in the NPU internal memory 120. Hereinafter, for convenience of description, the corresponding schedule may be referred to as the fourth schedule. The inference result value is sent to each component of the edge device 1000 to be utilized.

[0254] For example, when the inference result value is the result value of detecting a specific keyword, the neural processing unit 100 sends the inference result to the central processing unit 1080 so that the edge device 1000 can perform an operation corresponding to the specific keyword.

[0255] For example, the NPU scheduler 130 may drive the first to third processing elements PE1 to PE3 in the first schedule.

[0256] For example, the NPU scheduler 130 may drive the fourth to sixth processing elements PE4 to PE6 in the second schedule.

[0257] For example, the NPU scheduler 130 may drive the seventh and eighth processing elements PE7 and PE8 in the third schedule.

[0258] For example, the NPU scheduler 130 may output the inference result in the fourth schedule.

[0259] In summary, the NPU scheduler 130 may control the NPU internal memory 120 and the processing element array 110 to perform operations in the order of the first schedule, the second schedule, the third schedule, and the fourth schedule. That is, the NPU scheduler 130 may be configured to control the NPU internal memory 120 and the processing element array 110 to perform operations according to the set schedule order.

[0260] In summary, the neural processing unit 100 according to an example of the present disclosure may be configured to schedule the processing order based on the structure of the layer of the artificial neural network and the operation order data corresponding to the structure. The processing order to be scheduled may be at least one. For example, the neural processing unit 100 may predict all the operation orders, so that subsequent operations can be scheduled, or the operations can be scheduled in a specific order.

[0261] The NPU scheduler 130 may improve the memory reuse rate by controlling the NPU internal memory 120 by utilizing the schedule order based on the data locality information of the artificial neural network model or the information about its structure.

[0262] According to the characteristics of the artificial neural network operation driven in the neural processing unit 100 according to an example of the present disclosure, the operation value of one layer may be used as the input data of the subsequent layer.

[0263] Therefore, when the neural processing unit 100 controls the NPU internal memory 120 according to the scheduling order, the memory reuse rate of the NPU internal memory 120 can be improved.

[0264] Specifically, when the NPU scheduler 130 is configured to be provided with data locality information of the artificial neural network model or information about its structure, and calculates the order of operations for executing the artificial neural network through the provided data locality information of the artificial neural network model or information about its structure, the NPU scheduler 130 can identify the operation results of the node data of a specific layer of the artificial neural network model and the weight data of a specific connection network as the node data of the corresponding subsequent layer. Therefore, the NPU scheduler 130 can reuse the value of the memory address storing the corresponding operation result for subsequent operations.

[0265] For example, set the first operation value of the first scheduling as the node data of the second layer of the second scheduling. More specifically, the NPU scheduler 130 can reset the memory address value corresponding to the first operation value of the first scheduling stored in the NPU internal memory 120 to the memory address value corresponding to the node data of the second layer of the second scheduling. That is to say, the memory address value can be reused. Therefore, the NPU scheduler 130 reuses the memory address value of the first scheduling, so that the NPU internal memory 120 can utilize this data as the node data of the second layer of the second scheduling without a separate memory write operation.

[0266] For example, set the second operation value of the second scheduling as the node data of the third layer of the third scheduling. More specifically, the NPU scheduler 130 can reset the memory address value corresponding to the second operation value of the second scheduling stored in the NPU internal memory 120 to the memory address value corresponding to the node data of the third layer of the third scheduling. That is to say, the memory address value can be reused. Therefore, the NPU scheduler 130 reuses the memory address value of the second scheduling, so that the NPU internal memory 120 can utilize this data as the node data of the third layer of the third scheduling without a separate memory write operation.

[0267] For example, set the third operation value of the third scheduling as the node data of the fourth layer of the fourth scheduling. More specifically, the NPU scheduler 130 can reset the memory address value corresponding to the third operation value of the third scheduling stored in the NPU internal memory 120 to the memory address value corresponding to the node data of the fourth layer of the fourth scheduling. That is to say, the memory address value can be reused. Therefore, the NPU scheduler 130 reuses the memory address value of the third scheduling, so that the NPU internal memory 120 can utilize this data as the node data of the fourth layer of the fourth scheduling without a separate memory write operation.

[0268] In addition, the NPU scheduler 130 may be configured to determine whether to reuse the scheduling order and memory to control the NPU internal memory 120. In this case, the NPU scheduler 130 analyzes the data locality information of the artificial neural network model or information about its structure to provide an optimized schedule. In addition, the data required for the operations that can reuse the memory is not repeatedly stored in the NPU internal memory 120, thereby reducing the memory usage. In addition, the NPU scheduler 130 may optimize the NPU internal memory 120 by calculating to reduce the memory usage as much as the memory is reused.

[0269] The neural processing unit 100 according to an example of the present disclosure may be configured such that a variable value is input to an (N)-bit input that is the first input of the first processing element PE1, and a constant value is input to an (M)-bit input that is the second input. In addition, this configuration may be set in other processing elements of the processing element array 110 in the same manner. That is, one input of the processing element may be configured to receive a variable value, and the other input may be configured to receive a constant value. Therefore, the number of times of updating the data of the constant value can be reduced.

[0270] At this time, the NPU scheduler 130 uses the data locality information of the artificial neural network model 110a or information about its structure to set the input layer 110a-1, the first hidden layer 110a-3, the second hidden layer 110a-5, and the output layer 110a-7 as variables, and sets the weight data of the first connection network 110a-2, the weight data of the second connection network 110a-4, and the weight data of the third connection network 110a-6 as constants. That is, the NPU scheduler 130 can distinguish between constant values and variable values. However, the present disclosure is not limited to constant and variable data types, but distinguishes between frequently changing values and unchanging values to increase the reuse rate of the NPU internal memory 120.

[0271] That is, the NPU system memory 120 may be configured to save the weight data of the connection network stored in the NPU system memory 120 while maintaining the inference operation of the neural processing unit 100. Therefore, the memory read / write operations can be reduced.

[0272] That is, the NPU system memory 120 may be configured to reuse the MAC operation values stored in the NPU system memory 120 while maintaining the inference operation.

[0273] That is to say, the number of updates of the data storing the memory address for storing the input data (N bits) of the first input unit of the processing elements in the storage processing element array 110 can be greater than the number of updates of the data storing the memory address for storing the input data (M bits) of the second input unit. That is to say, the number of updates of the data of the second input unit can be less than the number of updates of the data of the first input unit.

[0274] Hereinafter, a convolutional neural network (CNN) as a type of deep neural network (DNN) in an artificial neural network will be mainly described.

[0275] A convolutional neural network can be a combination of one or more convolutional layers, pooling layers, and fully connected layers. A CNN has a structure suitable for two-dimensional data learning and inference and can be trained by a backpropagation algorithm.

[0276] Figure 5A The basic structure of a convolutional neural network is illustrated.

[0277] Reference Figure 5A , the input image can be represented by a two-dimensional matrix consisting of rows of a specific size and columns of a specific size. The input image can have multiple channels and the channels can indicate the number of color components of the input data image.

[0278] The convolution process refers to performing a convolution operation using a kernel while accessing the input image at a specified interval.

[0279] When a convolutional neural network moves from the current layer to the next layer, the weight values between the layers are reflected by the convolution to pass the weight values to the next layer.

[0280] For example, convolution is defined by two main parameters. The size of the input image (usually 1×1, 3×3, and 5×5 matrices) and the depth of the output feature map (the number of kernels) can be calculated by convolution. Convolution can start from a depth of 32, continue to a depth of 64, and end at a depth of 128 or 256.

[0281] Convolution can be operated by sliding a window of size 3×3 or 5×5 over a 3D input feature map, stopping at all positions and extracting 3D blocks of adjacent features.

[0282] The 3D block can be transformed into a 1D vector through the tensor product with the same learned weight matrix called weights. This vector can be spatially reassembled into a 3D output map. All spatial positions of the output feature map can correspond to the same position of the input feature map.

[0283] A convolutional neural network may include a convolutional layer that performs a convolution operation between a kernel (i.e., a weight matrix) that is trained through multiple gradient update iterations during the learning process and input data. If (m,n) is set as the kernel size and W is set as the weight value, the convolutional layer calculates the inner product to perform the convolution of the input data and the weight matrix.

[0284] The step size at which the kernel slides over the input data is called the stride, and the kernel region (m×n) can be called the receptive field. The same convolutional kernel is applied to different positions of the input, which may reduce the number of kernels to be learned. This also enables position-invariant learning. If there is an important pattern in the input, the convolutional filter can learn that pattern regardless of the position in the sequence.

[0285] The convolutional neural network can be adjusted or trained such that the input data is connected to a specific output estimate. The convolutional neural network can be adjusted using backpropagation based on the comparison between the output estimate and the ground truth until the output estimate gradually matches or approaches the ground truth.

[0286] The convolutional neural network can be trained by adjusting the weights between neurons based on the difference between the ground truth and the actual output.

[0287] Figure 5B The overall operation of the convolutional neural network is illustrated.

[0288] Reference Figure 5B , the input image is a two-dimensional matrix of size 5×5. Additionally, in Figure 5B , three nodes are used, namely Channel 1, Channel 2, and Channel 3.

[0289] First, the convolution operation of Layer 1 will be described.

[0290] The input image is convolved with Kernel 1 of Channel 1 at the first node of Layer 1, and the resulting output feature Figure 1 . Additionally, the input image is convolved with Kernel 2 of Channel 2 at the second node of Layer 1, and the resulting output feature Figure 2 . The input image is convolved with Kernel 3 of Channel 3 at the third node, and the resulting output feature Figure 3 .

[0291] Next, the pooling operation of Layer 2 will be described.

[0292] The features Figure 1 、features Figure 2 and features Figure 3Three nodes are input to Layer 2. Layer 2 receives the feature map output from Layer 1 as input to perform pooling. Pooling can reduce the size in the matrix or emphasize specific values in the matrix. Pooling methods can include max pooling, average pooling, and min pooling. Max pooling is used to collect the maximum value within a specific region of the matrix, and average pooling is used to calculate the average value within a specific region.

[0293] In Figure 5B the example of, the feature map of a 5×5 matrix is reduced to a 4×4 matrix through pooling.

[0294] Specifically, the first node of Layer 2 performs pooling with the features of Channel 1 Figure 1 as input and then outputs a 4×4 matrix. The second node of Layer 2 performs pooling with the features of Channel 2 Figure 2 as input and then outputs a 4×4 matrix. The third node of Layer 2 performs pooling with the features of Channel 3 Figure 3 as input and then outputs a 4×4 matrix.

[0295] Next, the convolution operation of Layer 3 will be described.

[0296] The first node of Layer 3 receives the output from the first node of Layer 2 as input to perform convolution with Kernel 4 and outputs the result. The second node of Layer 3 receives the output from the second node of Layer 2 as input to perform convolution with Kernel 5 of Channel 2 and outputs the result. Similarly, the third node of Layer 3 receives the output from the third node of Layer 2 as input to perform convolution with Kernel 6 of Channel 3 and outputs the result.

[0297] As described above, convolution and pooling are repeated. Finally, as Figure 5A shown, a fully connected layer can be output. The output can be input into the artificial neural network again for image recognition.

[0298] Hereinafter, the SoC will be mainly explained, but the disclosure of this specification is not limited to the SoC, and the content of this disclosure is also applicable to system-in-package (SIP) or printed circuit board (PCB) substrate-level systems. For example, each functional component is implemented by an independent semiconductor chip and is connected through a system bus, and the system bus is implemented by a conductive pattern formed on the PCB.

[0299] Figure 6 An exemplary architecture of a system-on-chip (SoC) including Figure 1 or 3 NPUs is shown.

[0300] Referring to Figure 6 , the exemplary SoC 100 includes multiple functional components, a system bus 500, a system in-circuit tester (ICT) 600, and multiple test wrappers 700a, 700b,..., and 700g, collectively referred to as test wrappers 700.

[0301] The multiple functional components may include an NPU core array 100-1, a central processing unit (CPU) core array 200, a graphics processing unit (GPU) core array 300, an internal memory 400, a memory controller 450, an input / output (I / O) interface 800, and a field programmable gate array (FPGA) 900.

[0302] Examples of the present disclosure are not limited thereto, and at least some of the multiple functional components may be removed. Examples of the present disclosure are not limited thereto, and may further include other functional components in addition to the above multiple functional components.

[0303] NPUs, CPUs, and GPUs are collectively referred to as general processing units (UPIs), application processing units (APUs), or application-specific processing units (ADPUs).

[0304] Each NPU core in the array 110-1 may refer to Figure 1 or the NPU100 in 3. In other words, in an array 110-a that includes Figure 1 or 3 of multiple NPUs 100 is illustrated in Figure 6 this.

[0305] Similarly, multiple CPU cores may be included in the array 200. Multiple GPU cores may be included in the array 300.

[0306] The array 100-1 of NPU cores may be connected to the system bus 500 via a wrapper 700a. Similarly, the array 200 of CPU cores may be connected to the system bus 500 through a wrapper 700b. Similarly, the array 300 of GPU cores may be connected to the system bus 500 via a wrapper 700c.

[0307] The internal memory 400 may be connected to the system bus 500 via a wrapper 700d. The internal memory 400 may be shared by CPU cores, GPU cores, and NPU cores.

[0308] The memory controller 400 connected to the external memory may be connected to the system bus 500 via a wrapper 700e.

[0309] The system bus 500 may be implemented by a conductive pattern formed on a semiconductor die. The system bus enables high-speed communication. For example, CPU cores, GPU cores, and NPU cores may read data from the internal memory 400 or write data to the internal memory 400 through the system bus 500. Further, CPU cores, GPU cores, and NPU cores may read data from the external memory or write data to the external memory through the memory controller 450.

[0310] The ICT600 can be connected to the system bus 500 through a dedicated signaling channel. In addition, the ICT600 can be connected to multiple wrappers 700 through a dedicated signaling channel.

[0311] Each wrapper 700 can be connected to the ICT 600 through a dedicated signaling channel. In addition, each wrapper 700 can be connected to the system bus 500 through a dedicated signaling channel. In addition, each wrapper 700 can be connected to the corresponding functional component in the SoC through a dedicated signaling channel.

[0312] For this purpose, each wrapper 700 can be designed to be located between the corresponding functional component in the SoC and the system bus 500.

[0313] For example, the first wrapper 700a can be connected to the array 100-1 of NPU cores, the system bus 500, and the ICT 600 through a dedicated signaling channel. The second wrapper 700a can be connected to the CPU core array 200, the system bus 500, and the ICT 600 through a dedicated signaling channel. The wrapper 700c can be connected to the GPU core array 300, the system bus 500, and the ICT 600 through a dedicated signaling channel. The fourth wrapper 700d can be connected to the internal memory 400, the system bus 500, and the ICT 600 through a dedicated signaling channel. The fifth wrapper 700e can be connected to the memory controller 450, the system bus 500, and the ICT 600 through a dedicated signaling channel. The sixth wrapper 700f can be connected to the I / O interface 800, the system bus 500, and the ICT600 through a dedicated signaling channel. The seventh wrapper 700g can be further connected to the I / O interface 800.

[0314] The ICT 600 can directly monitor the system bus 500 through each wrapper 700 or monitor the status of multiple functional components. Each functional component can be in an idle state or a busy state.

[0315] When a functional component in the idle state is found, the ICT 600 can select the corresponding functional component as the component under test (CUT).

[0316] If multiple functional components are in the idle state, the ICT 600 can select any one of the functional components as the CUT according to a predetermined rule.

[0317] If multiple functional components are in an idle state, the ICT 600 can randomly select any one of the functional components as the CUT. By doing so, the ICT 600 can disconnect or isolate the functional component selected as the CUT from the system bus 500. To this end, the ICT 600 can instruct the wrapper 700 connected to the functional component selected as the CUT to be disconnected or isolated. More specifically, the ICT 600 disconnects the functional component selected as the CUT from the system bus 500 through the wrapper 700, and then can instruct the wrapper 700 to send a signal to the system bus 500 instead of the functional component selected as the CUT. At this time, the signal transmitted to the system bus 500 can be the signal transmitted to the system bus 500 when the functional component selected as the CUT is in an idle state. To this end, when the functional component selected as the CUT is in an idle state, the wrapper 700 can monitor (or eavesdrop on) and store the signal transmitted to the system bus 500. The corresponding wrapper 700 regenerates the stored signal to transmit the regenerated signal to the system bus 500. At the same time, the corresponding wrapper 700 can detect the signal from the system bus 500.

[0318] Thereafter, the ICT 600 can test the functional component selected as the CUT.

[0319] Specifically, the rules can include rules according to the priority of the tasks to be executed, the priority rules between functional components, rules according to the presence or absence of spare parts corresponding to the functional components, rules defined by the number of tests, and rules defined by the previous test results.

[0320] For example, when the priority rule of the task indicates that the operation of the GPU has a higher priority than the operation of the CPU, between the idle CPU and GPU, the GPU can be preferentially tested. When the priority rule between functional components indicates that the priority of the CPU is higher than that of the GPU, between the idle CPU and GPU, the CPU can be preferentially tested. According to the rule of the presence or absence of spare parts, when the GPU is triple-core and the CPU is six-core, the GPU with fewer cores can be preferentially tested. According to the rule defined by the number of tests, when the CPU is tested 3 times and the GPU is tested 5 times, the tested CPU can be preferentially tested. According to the rule defined by the previous test result, when an abnormality is found in the previous test result of the CPU and the previous test result of the GPU is normal, the CPU can be preferentially tested.

[0321] When a conflict occurs due to accessing the functional component selected as the CUT from the system bus 500 at the start of the test or during the test, the ICT 600 can detect the conflict.

[0322] If so, the ICT 600 may stop (interrupt) the test and drive a back-off timer for the conflict.

[0323] The ICT 600 may restore the connection of the functional component selected as the CUT to the system bus 500.

[0324] Meanwhile, when the back-off time of the back-off timer for the conflict expires, the ICT 600 may monitor whether the functional component enters the idle state again. If the functional component enters the idle state again, the ICT 600 may select the functional component as the CUT again.

[0325] If no conflict is detected, the ICT 600 may continue the test and analyze the test results when the test is completed.

[0326] The test can be used to verify whether there are defects in the components of the system during its manufacture, whether they are damaged or have been broken. The damage or breakage may be caused by fatigue stress or physical stress (such as heat or electromagnetic pulse (EMP)) due to repeated use.

[0327] Performing tests on the NPU will be described. As described below, there are two types of tests, including functional tests and scan tests.

[0328] First, when performing a functional test on the NPU, the ICT600 may input a predetermined ANN test model and test input to the NPU. When the NPU outputs the inference result of the test input using the input ANN test model, the ICT 600 compares the expected inference result with the inference result from the NPU to analyze whether the NPU is normal or defective. For example, when the ANN test model is a predetermined CNN and the test input is a simple test image, the NPU performs convolution and pooling on the test image using the ANN test model and outputs a fully connected layer.

[0329] Next, when performing a scan test on the NPU, as will be described below, the ICT 600 may thread the flip-flops in the NPU through a scan chain. The ICT 600 may import the test input into at least one flip-flop and obtain the test result from the operation of the combinational logic of the flip-flop to analyze whether the NPU is defective or normal at runtime.

[0330] The tests performed by the ICT 600 can be tests performed before the SoC in mass production in the factory comes out to determine the fair quality. According to the present disclosure, it should be noted that tests for determining fair quality can also be performed during the runtime of the SoC. That is to say, according to the known technology, it is only possible to perform tests for determining fair quality before the SoC leaves the factory. However, according to the present disclosure, by finding the functional components in the idle state from among the multiple functional components in the SoC for sequential testing, it is possible to perform a fair quality test on the SoC during runtime.

[0331] As a result of the test analysis, when it is determined that the corresponding functional component is normal, the ICT 600 returns the connection to the functional component to the system bus 500. Specifically, the ICT 600 can disconnect the connection between the wrapper 700 and the system bus 500 and restore the connection between the functional component and the system bus 500. More specifically, the ICT 600 can initialize the functional component to be connected to the system bus 500 and then instruct the wrapper 700 to stop the signals to be transmitted to the system bus 500.

[0332] However, if the test analysis result is determined to be defective, the ICT 600 can repeat the test multiple times.

[0333] When, as a result of multiple repeated tests, the functional component is determined to be defective, that is, when it is determined that there are defects, has been damaged or has been broken in the manufacturing process of the functional component in the SoC, the ICT 600 can deactivate the functional component.

[0334] Alternatively, when the error code included in the one-time test analysis result indicates that there are defects, has been damaged or has been broken in the manufacturing process of the functional component in the SoC, the ICT 600 can deactivate the functional component.

[0335] To deactivate the functional component, the ICT 600 can cut off or disconnect the connection of the functional component determined to be defective to isolate the functional component determined to be defective from the system bus 500. Alternatively, to deactivate the defective functional component, the ICT 600 can power off (shut down) the functional component. When the functional component is powered off, the defective functional component can be prevented from malfunctioning, and the power consumption of the SoC can be reduced.

[0336] In addition, to deactivate the defective functional component, the ICT 600 can revoke the address of the functional component on the system bus 500 or send a signal for deletion to the system bus 500. That is to say, the ICT 600 can send a signal for deleting the address of the defective functional component to the component having the address used on the system bus 500.

[0337] Meanwhile, when the deactivation is completed, the ICT 600 can determine whether there is a spare part for the functional component. Even if a spare part may exist, when the spare part is not active, the ICT 600 can also activate the spare part. That is to say, the ICT 600 can send a signal including a request for updating the address of the activated spare part in the table to the component having an address table used on the system bus 500.

[0338] When the address on the system bus 500 is not assigned to the spare part in the deactivated state, the ICT 600 can send a signal to the system bus 500 for reassigning the address of the defective functional component to the spare part.

[0339] After monitoring whether the spare part is in the idle state, the ICT 600 can perform a test.

[0340] When there is no space for the deactivated functional component, the ICT 600 can allow the FPGA 900 to be programmed to mimic the same operation as the deactivated functional component. The information for programming the FPGA 900 can be stored in the internal memory 400. Alternatively, the information for programming the FPGA 900 can be stored in the cache memory of the FPGA 900.

[0341] As described above, when the FPGA 900 is programmed to mimic the same operation as the deactivated functional component, the ICT 600 can send a signal including a request for updating the address table used in the system bus 500. As an alternative, a signal including a request for reassigning the address of the faulty functional component to the FPGA can be sent to the system bus 500. In other words, the existing address of the FPGA can be revoked and replaced with the address of the faulty functional component.

[0342] When at least one functional component is determined to be defective, the SoC can be configured to display a warning message on a display device communicable with the SoC.

[0343] When at least one functional component is determined to be defective, the SoC can be configured to send a warning message to a server communicable with the SoC. Here, the server can be the manufacturer's server or the server of the service center. As described above, according to the present disclosure, the ICT and the wrapper are combined in the SoC and can perform tests during the runtime of the SoC.

[0344] Hereinafter, in order to understand the above content more deeply, a more detailed description will be given in conjunction with the table of contents.

[0345] I. Why Runtime Testing is Important

[0346] To prevent potential accidents that may be caused by hardware defects in autonomous systems, various studies have been conducted.

[0347] In various tests, including pre-deployment tests. According to this test technique, all hardware designs are checked before the product is sold to customers. After manufacturing is completed, the design is tested from various angles to detect and correct various problems that may be found during actual operation. For example, to test a chip design, test patterns are provided to perform scans of the inputs and checks of the output results. Although this technique can minimize potential problems in the hardware design before product shipment, defect problems during operation that may be caused by the aging of integrated circuits (ICs), the external environment, and the vulnerability of complex designs may not be resolved.

[0348] As described above, the above pre-deployment tests cannot effectively solve hardware defects, which makes the inventors become interested in testing methods during runtime.

[0349] From the perspective of the test mechanism, pre-deployment tests and post-deployment tests may seem similar, but there are obvious differences in when the tests can be conducted. Specifically, pre-deployment tests can only be conducted at specific times and are usually only allowed shortly after manufacturing. In contrast, tests during runtime can be conducted at any time under normal operating conditions.

[0350] There are two test techniques for tests during runtime, including functional tests and scan tests.

[0351] According to functional tests, test inputs are generated, and the output results obtained by inputting the generated test inputs into the original design are compared with the expected patterns. Alternatively, based on the original design, according to functional tests, input and output signals are monitored to detect abnormalities.

[0352] According to scan tests, an architecture for scan tests is inserted into the original design, and as many various test patterns as possible need to be created. As described above, after the scan architecture and test patterns are prepared, tests during runtime can be performed in various ways.

[0353] To perform scan tests, an ICT can connect multiple flip-flops in each CUT, import test inputs into at least one flip-flop, and obtain test results from the operations of the combinational logic of the flip-flops to analyze whether the CUT is defective or normal during runtime.

[0354] Figure 7 An example of a scan flip-flop is illustrated.

[0355] To more easily design hardware and minimize manufacturing defects, it is very important to apply design for testability (DFT).

[0356] Therefore, the architecture for scan testing is reflected in the design, and a test scope with a specific ratio of all detectable defects is defined to perform the test.

[0357] When using D-type flip-flops, the architecture for scan testing can be easily reflected in the design. During testing, all flip-flops in the CUT can operate as scan flip-flops, which include D flip-flops and multiplexers.

[0358] Compared with a normal D-type flip-flop as Figure 7 shown, the flip-flop can use two additional pins, namely the scan enable (SE) pin and the scan input (SI) pin. The SI pin is used for test input, and the SE pin can switch between the input for normal operation (D pin) and the test input for test operation (SI).

[0359] Figure 8 The figure illustrates an example of adding the architecture for scan testing in a hardware design.

[0360] As Figure 8 shown, all SE pins in the scan flip-flops are connected to the scan_enable (SE) port, the SI pin of each flip-flop is connected to the Q pin of the previous flip-flop or the scan input port, and the Q pin of each flip-flop is connected to the SI pin of the subsequent flip-flop.

[0361] These connections create multiple scan chains. That is, the flip-flops are interconnected to create scan chains.

[0362] When the SE (scan_enable) port is enabled, all scan flip-flops transfer data from the SI pin to the Q pin through the flip-flops, so that data can be transferred from the scan_in port to the corresponding scan_out port. All flip-flops on each scan chain transfer the test input from the scan_in port to the scan_out port.

[0363] The fewer the number of flip-flops on a scan chain, the faster the data can be shifted. However, the number of flip-flops on each scan chain and the number of scan chains are interdependent. The more scan chains are created, the fewer flip-flops there are on each scan chain (smaller).

[0364] II. Testing via ICT

[0365] The above tests are performed as background tasks so that the tests can be carried out without degrading the system performance. Based on the monitoring of the operations of the components to be tested, the ICT can determine whether the component is in an idle state. When the component is in an idle state, the test is performed so that the system performance is not degraded. The ICT continuously monitors the operating state of the CUT on the system bus, and the CUT may respond to unexpected accesses. When there is an access to the CUT, the operation of the CUT switches from the test operation to the normal operation to resume the CUT and return the CUT to the normal operation. The switching may have a slight time delay. According to the present disclosure, the system bus can be effectively used during the time delay to minimize the degradation of the system performance due to the resume.

[0366] II-1. Increase in the Complexity of the SoC Architecture

[0367] The design of integrated circuits (ICs) is becoming increasingly complex, and the integration level is also significantly increasing. The SoC is a highly integrated semiconductor device, such that defects in certain functional components may cause a degradation in the performance of the entire system. Therefore, it becomes increasingly important to perform tests to find defects in the functional components of the SoC.

[0368] From an operational perspective, Figure 9A illustrates Figure 6 the SoC.

[0369] The functional components (or IPs) can be divided into three types: 1) internal processors, 2) interface or communication controllers, and 3) memories. FIG. 9 shows the functional components (or IPs) 100 / 200 / 300, the internal memory 400, the memory controller 450, the system bus 500, the ICT 600, multiple wrappers 700a, 700b, 700c, 700d, 700f, and 700g (collectively referred to as 700), and the I / O interface 800.

[0370] The functional components (or IPs) 100 / 200 / 300 can perform functions related to encoding, decoding, encryption, decryption, and computing. The functional components (or IPs) 100 / 200 / 300 obtain the original data from the internal memory 400 and process the data with a specific algorithm. When the processing is completed, the output data can be sent together with a signal indicating completion.

[0371] The internal memory 400 can be a read-only memory (ROM) or a random access memory (RAM). The ROM corresponds to non-volatile memory, and the RAM corresponds to volatile memory.

[0372] A volatile memory is a type of memory that stores data only when powered and loses the stored data when power is cut off. Volatile memory can include static random access memory (SRAM) and dynamic random access memory (DRAM).

[0373] The internal memory 400 can include a solid state drive (SSD), flash memory, magnetic random access memory (MRAM), phase change RAM (PRAM), ferroelectric RAM (FeRAM), a hard disk, or flash memory. The internal memory 400 can also include synchronous random access memory (SRAM) and dynamic random access memory (DRAM).

[0374] The I / O interface can support various protocols and functions to enable the SoC to communicate with various external hardware.

[0375] However, when the SoC is built into an autonomous system to be used, the SoC may be damaged due to the aging of electronic components (e.g., transistors), physical effects, or the usage environment. Specifically, when a functional component (or IP) in the SoC processes important data, the damaged functional component may generate incorrect output data, which may significantly reduce the accuracy of the autonomous system.

[0376] To prevent this problem, as Figure 9A shown, the ICT 600 monitors the system bus 500 and monitors the states of the functional components 100 / 200 / 300, the internal memory 400, the memory controller 450, and the I / O interface 800 via the wrapper 700 or the system bus 500. When a functional component in the idle state is found, the ICT 600 selects this functional component as the CUT to perform a test.

[0377] In Figure 9A , during normal system operation, the connection to the system bus is represented by a dashed line, and the signals of the ICT and the wrapper are represented by solid lines.

[0378] Figure 9B Illustrated is the configuration for testing the NPU.

[0379] Refer to Figure 9B , the NPU 100 can also include another component for testing the NPU. Specifically refer to Figure 9B , at least one of a random number generator, a predetermined test data memory, and a temporary register can be selectively further included in Figure 1 or the NPU100 and components shown in 3. A MUX can be set between the NPU internal memory 120 and the processing element array 110 to perform an internal test of the NPU 100. The MUX can be configured to switch components configured to test the processing element array 110 and the NPU internal memory 120.

[0380] A method for testing the processing element array 110 using random numbers will be described. Figure 9B The random number generator in the illustrated NPU 100 can generate random numbers based on a predetermined seed. The MUX selects at least one of the processing element arrays 110 to test whether the NPU 100 is defective.

[0381] The ICT 600 monitors the state of the NPU 100 via the wrapper 700a, and when it is determined that the NPU 100 is in an idle state, the ICT 600 can command the NPU 100 to start the test.

[0382] By way of a specific example, the ICT 600 selects at least one of the multiple PEs included in the NPU 100 to command the start of the test.

[0383] By way of a specific example, when the ICT 600 determines that a predetermined percentage of the PEs (e.g., 20% of all PEs) included in the NPU 100 are in an idle state, the ICT 600 can command the NPU 100 to start the test. In other words, when the ratio of idle PEs among all PEs is greater than or equal to a threshold value, the ICT can command the start of the test.

[0384] By way of a specific example, when the ICT 600 selects a predetermined percentage of the PEs (e.g., 50% of all PEs) included in the NPU 100 to command the NPU 100 to start the test.

[0385] When testing the NPU 100, the inference speed of the NPU, i.e., the inferences per second (IPS), may decrease. Specifically, the inferences per second can be reduced according to the number of PEs to be tested. By way of a specific example, when 50% of all PEs are tested, the inferences per second may be reduced by approximately 50%, and when 30% of all PEs are tested, the inferences per second may be reduced by approximately 30%.

[0386] Therefore, according to the example, the NPU 100 may further include additional PEs to improve the speed reduction caused by the test.

[0387] As another example, when the NPU 100 operates at a value below a predetermined IPS value, the ICT 600 may instruct the NPU 100 to perform a test. Specifically, assuming that the NPU 100 operates at a maximum of 100 IPS and the IPS threshold is 30 IPS, if the NPU 100 operates at 30 IPS or higher, the ICT 600 may instruct the NPU 100 to perform a test during the remaining time. For example, when the NPU 100 operates at 40 IPS, the remaining 60 IPS can be used for testing. Therefore, a significant speed reduction of the NPU is not caused.

[0388] As another example, when the data sent to the NPU internal memory 120 in the memory 400 is delayed such that the NPU 100 is in an idle state or enters a data shortage period, the ICT 600 may instruct the NPU 100 to perform a test.

[0389] When a test is performed on the NPU 100, the register file RF corresponding to each PE in the NPU 100 is initialized with predetermined test input data, and the corresponding PE can perform inference based on the test input data in the register file RF. The predetermined test input data can be a functional test or a partial functional test of the NPU.

[0390] When the NPU 100 is tested, as described above, the random number generator in the NPU 100 generates a random number. By doing so, the register file RF is initialized with the generated random number, and the corresponding PE performs inference based on the random number in the register file RF.

[0391] Alternatively, the ICT 600 commands the CPU 200 via the wrapper 700b to import the test input data into the register file RF in the NPU 100.

[0392] When the NPU 100 is tested, multiple register files RF in the NPU 100 are initialized with a single test input data, and the corresponding PE can perform inference based on the test input data in the register file RF. Specifically, multiple PEs in the NPU 100 can be tested based on the same single test input data, and the inference results are output.

[0393] Alternatively, when the NPU 100 is tested, some of the register files RF in the NPU 100 are initialized based on specific test input data, and the corresponding PE can perform inference based on the test input data in the register file RF.

[0394] The register file RF can reset the flip-flops in each PE and transfer the test input data to the PE as described above.

[0395] For example, the size of each RF can be 1Kb.

[0396] II-2. Necessity of Packaging

[0397] Figure 10 The operation of the wrapper is illustrated.

[0398] As described above, ICT can test many functional components (i.e., IP, I / O interfaces, memories, etc.) in the SoC during the runtime of the SoC. For this purpose, during the test of the functional component selected as the CUT, it is necessary to solve the conflict problem caused by accessing the functional component from the system bus.

[0399] To solve the conflict problem, after monitoring whether the functional component is in the idle state, when it is detected that the functional component is in the idle state, the functional component is switched from the normal operation mode to the test operation mode, and then the test needs to be executed. When a conflict is detected during the test, the functional component needs to be switched to the normal operation mode. After switching the operation to the normal operation mode, the functional component needs to correctly process the input data.

[0400] For this purpose, the shown wrapper 700 needs to be set between the functional component and the system bus 500. The wrapper 700 can include multiplexer gates that selectively control the input and output of each operation mode.

[0401] As shown in the figure, when the TEST_ENABLE port is open, the test vector can be input to the CUT, and the TEST_OUTPUT port can send the output. The general data output from the wrapper 700 can be transmitted to other functional components through the system bus. In contrast, the test result can be directly transmitted to the ICT 600. The ICT 600 can receive the test vector for testing from an external memory or an internal memory, and store the test result in the internal memory or the external memory, or transmit the test result to the outside.

[0402] To test the SoC at runtime, the ICT 600 can perform multiple processes. First, the ICT 600 can select a functional component to be tested as the CUT based on a predetermined rule. Since the CUT needs to respond to access from the system bus when the SoC is operating, it is effective to select as many functional components in the idle state as possible as the CUT. For this purpose, the ICT 600 can monitor whether the functional component enters the idle state. When the functional component enters the idle state, the wrapper 700 can open the TEST_ENABLE port. The ICT600 can import the test vector into the CUT through the TEST_ENABLE port.

[0403] The ICT 600 can collect and analyze test results from the CUT via the TEST_OUTPUT port of the wrapper 700. When the test results indicate that a problem has been detected, the ICT 600 can perform post actions. During testing, when a general access to the CUT from the system bus 500 is detected, the ICT 600 can temporarily delay the access from the system bus 500 and then immediately stop (interrupt) the test operation. Thereafter, the ICT 600 can restore the previous values of the register settings for the CUT and turn off the TEST_ENABLE port of the wrapper 700. When the normal operation of the CUT is ready, the ICT 600 can control the wrapper 700 to return the connections for input and output to the CUT to the system bus 500.

[0404] Figure 11 The figure illustrates the internal configuration of the ICT.

[0405] Reference Figure 11 , the ICT 600 can include a configuration data (CONF_DATA) restorer 610, a state detector 620, a scheduler 630, a tester 640, a test vector generator 650, a host interface 660, and a post-action (POST_ACT) unit 670.

[0406] The state detector 620 can detect whether a functional component in the SoC chip is in an idle state or a busy state (or a processing state). When any functional component enters the idle state, the state detector 620 sends the ID (C_ID) of the functional component to the scheduler 630 to perform a test.

[0407] The scheduler 630 can manage the overall operation of the ICT 600. The scheduler 630 can receive the state of the functional component from the state detector 620 and trigger a test. The scheduler 630 can send the ID of the component to the tester.

[0408] The tester 640 controls the wrapper 700, transmits test vectors, obtains test results, and then compares whether the test results match the expected test results. Thereafter, the tester 640 can transmit the test results to the post-action unit 670. The tester 640 can restore the register settings of the functional component selected as the CUT to their original values.

[0409] The test vector generator 650 can generate test vectors (or predefined test input data) and corresponding expected test results. The test vector generator 650 can include a buffer, a memory interface, a memory for storing test vectors and expected test results, and a random number generator. When the test starts, a test pattern for generating test vectors can be loaded into the buffer. The random number generator can be used to generate test vectors. The random number generator can cause the memory not to store all test vectors but to generate various test vectors.

[0410] When receiving the ID (e.g., C_ID) of the defective functional component found from the tester 640, the post-action unit 670 can perform post-actions. The post-actions can isolate the defective functional component or notify the defect to the user or the remote host device.

[0411] The host interface 660 can report the defective functional component found during the test to the user or the remote host device. If there are changes related to the test operation, the host interface 660 can notify the remote host device.

[0412] When the test is completed or an access to the functional component selected as the CUT from the system bus 500 is detected during the test, the configuration data restorer 610 can restore the register settings of the CUT so that the tester 640 can switch the CUT to the normal operation mode. Most functional components will have specific register setting values for normal operation. Accordingly, the configuration data restorer 610 can store the register setting values of the functional component before performing the test and restore the register setting values to the functional component when it is necessary to switch the CUT to the normal operation mode.

[0413] II-3. Detecting the idle state of the functional component

[0414] Figure 12 The figure illustrates the operation of monitoring whether the functional component is in the idle state through ICT.

[0415] To detect whether a functional component is in an idle state during normal operation mode, the ICT 600 can use one or both of two techniques. First, the ICT 600 can monitor whether the component is in an idle state or in use based on some hardware signals that directly or indirectly indicate operation. For example, the ICT 600 can monitor the power gating control signal to disconnect the functional component, thereby reducing the power consumption of the functional component. In addition, the ICT 600 can determine whether the functional component is in an idle state based on an output signal that directly or indirectly indicates whether the component is operating or the value of a register that stores information related to the operation in the functional component. Second, the ICT 600 monitors signals from the system bus via the wrapper 700 or monitors the input / output ports of the functional component during a specific time period to determine whether the functional component is in an idle state.

[0416] II-4. Handling of access conflicts

[0417] Figure 13 Illustrates the operation among a host, a slave, and an arbiter operating on the system bus.

[0418] The host on the system bus can be an entity that uses the slave, the slave can be an entity used by the host, and the arbiter can be an entity that arbitrates and makes decisions between the host and the slave.

[0419] Figure 13 The slave shown in can be a functional component selected as the CUT and the arbiter can be the ICT.

[0420] When detecting an access to normal operation from the system bus 500 while testing the functional component selected as the CUT, the ICT 600 may require a predetermined amount of time or more to restore the CUT to its previous state. The ICT 600 can temporarily deactivate (or de-assert) the HREADY signal to temporarily stop system access from the host, stop (interrupt) the test activity, restore the register settings of the CUT, and change the direction of the data input to or output from the wrapper. When the CUT as the slave is ready to perform a task with the host, the HREADY signal can be turned on. However, according to the present disclosure, the ICT may cause some time delay in the bus separation operation. The specific process will be described below.

[0421] First, the host activates (or asserts) the HBSREQ signal for bus access. Second, during the arbitration or determination process, the arbiter activates (or asserts) the HGRANT signal to allow bus access. By doing so, the host can transfer data to the CUT acting as a slave via the system bus. If the ICT is performing a test processing operation, the ICT sends the HSPLIT signal along with the bit indicating the current host to the arbiter, and at the same time activates (or asserts) the SPLIT signal in the HRESP signal. After activation (assertion), the host cancels the access to the CUT, and the arbiter performs the arbitration or determination process without host intervention. When the CUT is ready to respond to the access from the host, the ICT deactivates the HSPLIT signal, and the host waits for authorization from the arbiter to resume the task of accessing the CUT.

[0422] Figure 14 An example of adding a shift register in the SoC chip is shown.

[0423] The inventors of the present disclosure have recognized that access to the I / O interface may not cause a conflict on the system bus. For example, when the target CUT is the host, the external device connected via the I / O interface does not request access for itself, so no conflict occurs. Therefore, it may be effective to only focus on solving the conflict problem that occurs when the CUT is a slave.

[0424] Instead, in order to delay the data transmitted from the external device to the CUT during the recovery time, a shift register can be added between the port of the SoC and the external interface port of the CUT.

[0425] A shift register can be added to store the access signal input from outside the SoC while recovering the CUT. When the CUT is ready, the access signal is regenerated by the shift register for output.

[0426] The depth of the shift register can be determined by the number of clock cycles required to recover the CUT to normal operation. Specifically, when one or more functional components need to receive signals from outside the SoC, the depth of the shift register can be variable. In this case, the depth of the shift register can be determined by the ICT.

[0427] II-5. Operation Sequence of ICT

[0428] Figure 15 The operation sequence of the ICT is illustrated.

[0429] Reference Figure 15 , when the timer related to the start of the ICT test during runtime expires (S601), the ICT monitors whether any functional component is in an idle state and detects the functional component in an idle state (S603).

[0430] By doing so, the ICT performs a test preparation process (S605). The test preparation process may include selecting a functional component as the CUT, isolating the functional component selected as the CUT from the system bus, and generating test vectors as test input data. The isolation from the system bus may mean that the ICT changes the directions of the inputs and outputs on the wrapper that communicates with the functional component selected as the CUT.

[0431] The ICT imports the test vectors as test input data into the CUT (S607).

[0432] When the test is completed normally, the ICT checks the test result (S609). For the check, the ICT may compare whether the test result matches the expected test result.

[0433] When the test result indicates that the functional component selected as the CUT has no problems (that is, no defects or damages), the ICT may restore the functional component to the normal operating state (S611).

[0434] Meanwhile, when an access to the functional component selected as the CUT is detected from the system bus during test preparation or testing, the ICT may restore the functional component selected as the CUT to the normal operating state (S613). The restoration may mean that the register setting values of the functional component selected as the CUT are restored and the directions of the inputs and outputs return to the original state on the wrapper that communicates with the functional component selected as the CUT.

[0435] In this case, the ICT drives a back-off timer (S615) and when the back-off timer expires, may return to step S603.

[0436] Meanwhile, when the test result indicates that the functional component selected as the CUT has problems (that is, defects or damages), the ICT may perform a post-operation (S617).

[0437] II-6. Testing of the internal memory

[0438] Figure 16 The figure illustrates the test process of the internal memory.

[0439] The testing of the internal memory may be different from the testing of the functional components. Two testing techniques for the internal memory will be presented below.

[0440] The first technique is a technique that uses an error detection code to detect errors during the process of reading data from the internal memory. If the error detection code obtained during the reading process is different from the predetermined error detection code, the ICT may determine the code as an error.

[0441] The second technique is a technique that performs read and write tests in a hard manner during normal operation.

[0442] Figure 16 The second technique is illustrated. The test logic for encapsulating the internal memory can perform read and write tests during system operation and bypass access from the system bus. To fully process the test, the tester in the ICT can be responsible for address management. The illustrated temporary register file can temporarily store the original data that is easily deleted due to the test. After the test is completed, the original data in the temporary register file is recorded in the internal memory again.

[0443] If an unpredictable access occurs during the test, the data on the system bus can be recorded in the temporary register file. Conversely, the data in the temporary register file can be moved to the system bus.

[0444] The test technique described above can be applied not only to the internal memory but also to the external memory in the same way.

[0445] II-7. Operations after Testing

[0446] When there are hardware defects in the SoC, the operations after testing can be very important. For example, notifying the user of the defect to recommend stopping use. For this purpose, Figure 11 the post-action unit 670 can provide information about the detected defective functional component and information about the test input data (that is, the test vector) that caused the defect. The above information can enable the user to know the location of the defective functional component. It is necessary to stop and isolate the use of the detected defective functional component. To prevent the defective functional component from degrading the performance of the entire system, the output signal of the functional component can be replaced with a predetermined signal. Alternatively, the functional component can be reset or gated. Alternatively, power gating can be performed on the functional component.

[0447] At the same time, when the functional component is isolated, the SoC faces another problem. Therefore, even if some functional components are defective, a method that allows the SoC to still operate needs to be proposed. For example, when the SoC is installed in a product that requires high reliability, the SoC also needs to include spare parts for some functional components. If some functional components are defective, the spare parts may operate in place of the functional components. However, when some functional components are replicated, it may increase the area of the semiconductor device. To solve this problem, adding programmable logic to the SoC may be effective.

[0448] III. Functional Testing or Testing of Functional Combinations during SoC Runtime

[0449] Figure 17 The process of testing a function using a random number generator is illustrated.

[0450] A functional test is a test that imports test input data (e.g., test vectors) into the CUT and compares whether the output from the CUT matches the expected output. For a correct evaluation based on the comparison, each input data needs to accurately derive the expected output. The test scope of the test input data needs to be large enough to detect all defects.

[0451] In a specific design, two test input data can be used for functional testing. First, a random number generator connected to an XOR operation can be used for Figure 17 the test operation shown. Generally, a random number generator can generate a pseudo-random number stream based on an input seed. The random number stream is imported into the CUT through a wrapper, and the output is accumulated through the XOR operation and stored in a test result register. When the test is completed, the value stored in the test result register can be compared with the expected result corresponding to the test input data. If there is a difference in the comparison result, an error notification may be issued.

[0452] Second, all test patterns of the test input data and the corresponding predicted results can be fixed respectively and stored in the internal memory of the SoC or in an external memory. When the test input data (i.e., test vectors) from the memory is input into the CUT, the output of the CUT can be compared with the expected result corresponding to the test input data.

[0453] To perform functional testing during the runtime of the SoC, ICT plays an important role in transmitting data, communicating with the system bus, monitoring the status of the CUT, etc. Specifically, when the CUT is in an idle state, ICT needs to determine when to perform the test. During the test, the random number generator generates a random number stream as test input data and transmits the test input data to the CUT. If there is a difference between the test result and the expected test result, ICT transmits the information to the post-operation unit.

[0454] During the functional test process, functional components in the SoC can be used, so generally, the frequency of test operations needs to be lower than or equal to the frequency of normal operations to avoid timing differences (i.e., timing violations). To perform tests in real time during normal operation, it is effective to perform tests when the functional components are in an idle state. Therefore, there is no alternative but to perform tests at a high frequency.

[0455] IV. Performing Tests during SoC Runtime Using a Combination of DFT (Discrete Fourier Transform) and ICT

[0456] IV-1. Multiple Clocks

[0457] Figure 18A An example of multiple clocks is illustrated. Figure 18BIt is an example diagram showing the operation of the tester at multiple clocks, and Figure 18C illustrates the path of the test input data.

[0458] During the test, there are two techniques for importing a test input data (i.e., test vector).

[0459] The first technique is to use a time period to "shift the data" as Figure 18A shown. The SE (scan enable) port is enabled, and the Q output of the flip-flop is connected to the D input of another flip-flop. This connection can form a scan chain that connects the scan input to the scan output through the flip-flop chain.

[0460] Therefore, all the combinational logic of the design can be disabled, and there may be no reference logic units for the data path (that is, the path from one flip-flop to another).

[0461] When T cycle is defined as the clock period, T launch is defined as the time delay from the clock source of the first flip-flop to the CK pin, T capture is defined as the time delay from the clock source to the CP pin of the second flip-flop, T clk2q is defined as the time delay from the CK to the Q pin of the first flip-flop, T dpmax is defined as the time delay from the Q of the first flip-flop to the D of the second flip-flop, T cycle >T launch +T clk2q +T dpmax +T setup +T margin –T capture .

[0462] When scan testing is enabled, from the perspective of scan testing, T dpmax can be reduced to zero. Ideally, T dpmax may be zero. However, to solve the timing violation, when adding multiple inverters or buffers, the time delay may be greater than zero.

[0463] Alternatively, T dpmax >>T clk2q +T setup +T launch –T capture . During the time period of "shifting the data", it will be processed at a higher frequency.

[0464] In as Figure 18ADuring the time period of "capturing data" shown, the scan enable pin is deactivated and thus the functional component is re-enabled, and combinational logic can be enabled on the data path. To address timing violations during data capture, a time delay can be added between the clock at one end of the "shifting data" time period and the clock at one end of the "capturing data" time period.

[0465] The delay between clock cycles can be greater than or equal to the clock cycle of normal operation. To detect when the "shifting data" time period is complete based on the maximum number of flip-flops on the scan chain corresponding to the shift value, a counter is added, and to manage the time delay within the "capturing data" time period, another counter can be added.

[0466] In Figure 18B the test block receives two input clocks. One is f_clk for normal operation, and the other is sclk for "shifting data". "Clock configuration" is inserted into the tester block so that the s_clk signal can be set to be used during the "shifting data" period and the "capturing data" period.

[0467] To control the switching between f_clk for normal operation and s_clk for test operation, the TE signal corresponding to the CUT can be used. When the ID of the component (i.e., C-ID) is received from the scheduler, the test block in the ICT is ready for testing. The TE of the CUT available through the decoder can enable the test process.

[0468] Figure 19A illustrates an example of a functional component, Figure 19B illustrates an example where test input data (e.g., test vectors) are imported into the tester in the ICT.

[0469] To apply the discrete Fourier transform (DFT) in testing during the runtime of the SoC, scan chains are added in the CUT, and all flip-flops can be surrounded by scan flip-flops. The scan input, scan output, and TEST_ENABLE and SCAN_ENABLE signaling are connected to the tester in the ICT, and the original inputs and original outputs of the CUT can communicate with the system bus through the tester and wrapper.

[0470] As Figure 19B shown, from the perspective of the memory storing the test mode, the block can be divided into four parts. The first part is the part for storing input shift vectors, the second part is the part for storing output shift vectors, the third part is the part for storing input capture vectors, and the fourth part is the part for storing output capture vectors. To start the test, the input shift data is loaded from the memory and input into the CUT through the tester.

[0471] In each scan chain, after all flip-flops are filled with a shift vector, when loading a first input capture vector including values of scan inputs and initial inputs, a first output capture vector including values of all scan outputs and initial outputs is loaded and then compared with actual output capture data. Each loaded shift vector is accompanied by output shift data, and actual output data can be compared with the output shift vector or the output capture vector.

[0472] Figure 20 FIG. illustrates a test process using DFT. Figure 21 FIG. illustrates an example of shift data and capture data during the test process.

[0473] During the step of data shifting, when the scan_enable port is enabled, the SCAN_IN port can be connected to the SCAN_OUT port through flip-flops without combinational logic. Input shift vectors can be loaded in all scan chains until all flip-flops have values shifted from the input shift vectors. One shifted value can pass through one flip-flop per clock cycle. That is, the D pin of the previous flip-flop can be connected to the D pin of the subsequent flip-flop.

[0474] When the scan_enable port is disabled during the capture step, the D pins of all flip-flops are not connected to the Q pins of the previous flip-flops but can be directly connected to combinational logic.

[0475] The capture vector output can be loaded into the Q outputs of all flip-flops at the positive (+) edge of the clock cycle through combinational logic. In the first data capture step, a data transfer process is prepared to compare the output data with the expected output data, and then the comparison is performed at each positive clock edge. After loading all test vector inputs, the process returns to the first data shifting step and each process starts over.

[0476] Figure 21 FIG. illustrates the shifting and capture processes. Figure 21 The rectangular boxes in represent flip-flops in each scan chain, and all flip-flops are filled at the end of the data shifting step.

[0477] Figure 22 FIG. illustrates an example of switching the test mode to the normal operation mode.

[0478] As is known in reference Figure 22 During the output test mode, the data shifting process and the capture step can be repeated. If there is access to the CUT, the CUT will return to the normal operation mode and the test will be aborted. Thereafter, a skip mode is executed during a predetermined time period, and then the output test mode can be executed again.

[0479] Figure 23 Illustrates an example of the operation of a flip - flop on a scan chain, and Figure 24 illustrates a part of the CUT operating in the normal operation mode.

[0480] When an unexpected access to the CUT from the system bus occurs, TEST_ENABLE is disabled, and data shifting or capture can be quickly stopped. The CUT resumes to the normal operation mode and the test can be reverted.

[0481] When the CUT enters the idle state again, the previous data shifting step can be restarted for testing. However, in the first transition step after the transition from the normal operation mode to the test operation mode, the comparison of the output results is disabled and the comparison of the output results can be performed from the subsequent capture step.

[0482] That is, as Figure 23 shown, the shifted input values are not loaded into all flip - flops of the scan chain, so the comparison can be skipped.

[0483] V, analog

[0484] To verify the above, simulations from simple sub - cases to complex sub - cases were performed using the electronic design automation tools of Synopsys, Inc. When the verification is successfully executed, the design implemented by software code can be converted into logic gates. As a next step, all D flip - flops (DFFs) can be replaced by scan flip - flops, and a scan chain can be generated. To increase the coverage by the test, the netlist can be repeatedly modified and tested. ATPG tools can be used to generate test vectors and expected results. When the test mode of the scan and the inserted netlist are ready, ICT can be applied to the design. The details of each test case will be described with reference to the accompanying drawings.

[0485] Figure 25 Illustrates the process for simulation.

[0486] The simulation process shown uses the design compiler tool of Synopsys, Inc. The Design Compiler (DC) can also be used to convert software code into logic gate level based on timing constraints (such as period, transition, capacitance, time information included in the library package). When all the constraints are met, value optimization can be repeated. When the design does not meet the requirements, the constraints can be adjusted.

[0487] The output of the DC tool including the above - mentioned timing constraints can be used as the input for the design for test (DFT).

[0488] In the scan test import step, the number of scan ports and the scan chains can be set. Usually, to minimize additional ports during design, the original input and output ports are used to generate scan ports. Additionally, the number of input scan ports is equal to the number of scan chains. The more the number of scan chains, the fewer the number of shift clock cycles for shifting data. Therefore, maximizing the number of scan chains is the best choice for testing. After the scan setup is complete, the DFT compiler can replace all flip-flops with scan-enabled flip-flops and connect the scan input (scan_in) pins and scan output (scan_out) pins to each other to generate scan chains. The additional connections and scan flip-flops make the design more complex and cause time delays in most of the data paths. Therefore, the DFT compiler can continuously optimize power and timing after connecting the scan chains. After completing the scan test import, the DFT DRC that follows the DFT rules checks if all test connections are connected. The test input data (i.e., test vectors) is ready.

[0489] To check the test coverage and generate test patterns, the output of the DFT compiler is input to Tetramax from Synopsys, Inc. When the test coverage does not meet the expected requirements, the scan test import step can be executed again to modify the design. This task can be repeated until the desired test coverage is obtained.

[0490] V-1. Design Experiment

[0491] As a design experiment, a JPEG image encoder can be used. This test can use approximately 265,118 combinational cells, 72,345 sequential cells, and 31,439 inverter / buffer cells. As a result of being executed based on cell library information and many thresholds, it is confirmed that the frequency that meets the timing limit is 100 MHz and the frequency for shifting test patterns is 1 GHz. Approximately 512 scan chains are used, and the maximum number of flip-flops used on each scan chain is 75. Therefore, it is confirmed that the time period for shifting test patterns is approximately 75 cycles, which corresponds to 75 ns. Capturing data requires one cycle, which corresponds to approximately 10 ns. Approximately 256 test patterns are input, each test pattern includes approximately 75 test vectors for shifting, and approximately one test vector can be used for capturing. To complete one test cycle, 13,260 (156×75 + 156×10) ns is required.

[0492] Figure 26 The test architecture of the JPEG image encoder is illustrated.

[0493] As described above, to check whether the ratio of controllable nodes and observable nodes is sufficient for testing, Tetramax is used to confirm the test range. Usually, a 99.97% test range and 1,483.5 K bytes are used for testing.

[0494] From the perspective of power consumption, each test is measured by different types of inputs. In Table 1 below, the internal power, switching power, and leakage power of each input are shown. Specifically, first, an input that cannot be controlled for the test mode and normal mode is implemented. This is called static power measurement. After inputting to the power compiler tool, to estimate the power consumption in the test mode, the TE (test_enable) port is turned on, and to estimate the power consumption in the normal operating mode, the TE (test_enable) port is turned off. Second, to estimate the power consumption, the input is controlled by a specific time interval. This is called dynamic power estimation. The control of the input can be divided into three modes. First, the test mode is turned on and the test mode is provided. Second, the test mode is turned off, and the input for normal operation is provided. Third, after completing the test of each test mode, the mode is switched to the normal operation mode. To obtain its power consumption value, the test mode and normal operation mode are switched.

[0495] The following table is the power consumption value of the JPEG image encoder.

[0496] Table 1

[0497]

[0498] Similar to the JPEG image encoder, the functional components of the Advanced Encryption Standard (AES) were tested. In addition, the functional components for image classification in autonomous vehicles were also tested. The results are shown below.

[0499] The following Table 2 represents the AES design.

[0500] Table 2

[0501] Number of combinational cells 160,261 Number of sequential cells 11,701 Number of buffers / inverters 22,377 Full area 400,464.797234 Frequency in normal operation mode 100Mhz Frequency in test mode 1Ghz Propagation time 46ns Capture time 10ns Number of test modes 315 Test range 100% Memory size 948.5KB

[0502] The following Table 3 represents the test of the functional components of AES.

[0503] The following Table 3

[0504]

[0505] The following Table 4 shows the details of CONVO2.

[0506] Table 4

[0507] Number of combinational cells 2,245,932 Number of sequential cells 424,695 Number of buffers / inverters 154,510 Frequency in normal operation mode 50Mhz Frequency in test mode 1Ghz Propagation time 829ns Capture time 20ns Number of test modes 183 Test range 100% Memory size 18.634MB

[0508] The following Table 5 shows the power consumption of CONVO2.

[0509] Table 5

[0510]

[0511] Functional testing and testing imported by scanning each have their own advantages and disadvantages. The testing imported by scanning has the disadvantages of using more memory and time delay compared with functional testing, and has the advantage of a wide testing range.

[0512] Specifically, when the SoC is installed in a product that requires high reliability such as an autonomous driving vehicle, the scan-import type testing with a wide testing range would be advantageous. In addition, the scan-import type testing can increase the frequency of test operations and reduce the test time. When a long test time is required, it may increase the possibility of a car accident, making it undesirable. The scan-import type testing increases the frequency of test operations, so that more test patterns can be imported during idle time, and hardware defects in the SoC can be detected faster. The advantage of ordinary functional testing is low power consumption, but in an environment that requires high reliability, such as an autonomous driving vehicle, power consumption is not important.

[0513] SoC has been mainly explained so far, but the disclosure of this specification is not limited to SoC, and the content of this disclosure is also applicable to system-in-package (SIP) or printed circuit board (PCB)-based board-level systems. For example, each functional component is implemented by an independent semiconductor chip and connected through a system bus, and the system bus is implemented by a conductive pattern formed on the PCB.

[0514] The embodiments of the present disclosure disclosed in this specification and the drawings only provide a specific example for easy description and better understanding of the technical description of the present disclosure, and are not used to limit the scope of the present disclosure. It is obvious to those skilled in the art that other modifications are possible in addition to the examples described so far.

[0515]

National R & D Project Supporting the Present Invention

[0516]

Project Identification Number

[0517]

Task Number

[0518]

Department Name

[0519]

Name of the Task Management (Professional) Institution

[0520]

Research Project Name

[0521]

Research Task Name

[0522] System

[0523]

Contribution Rate

[0524]

Organization Name Executing the Task

[0525]

Research Period

Claims

1. A system-on-chip (SoC) for testing components in a system during runtime, the system-on-chip comprises: a plurality of functional components, each of the plurality of functional components including circuitry; a system bus configured to allow the plurality of functional components to communicate with each other; one or more wrappers, each of the one or more wrappers being connected to one of the plurality of functional components; and an in-system component tester (ICT) configured to: monitor the states of the plurality of functional components via the one or more wrappers; select at least one functional component among the plurality of functional components that is in an idle state as a component under test (CUT); test the at least one functional component selected as the component under test through the one or more wrappers; interrupt a test step regarding the at least one functional component selected as the component under test based on detecting a conflict with an access from the system bus to the at least one functional component selected as the component under test; and allow the at least one functional component to be connected to the system bus based on the interrupt step.

2. The system-on-chip according to claim 1, wherein the in-system component tester is further configured to: after allowing the at least one functional component to be connected to the system bus, if as a result of the monitoring step, the at least one functional component is again in the idle state, return to the selection step.

3. The system-on-chip according to claim 2, wherein the return to the selection step occurs after an expiration of a backoff time regarding the conflict.

4. The system-on-chip according to claim 1, wherein the plurality of functional components includes one or more universal processing units (UPUs).

5. The system-on-chip according to claim 4, wherein the one or more universal processing units include at least one of the following: one or more central processing units (CPUs); one or more graphics processing units (GPUs); and one or more neural processing units (NPUs) configured to perform operations of an artificial neural network (ANN) model.

6. The system-on-chip according to claim 4, wherein the plurality of functional components further includes at least one of the following: at least one memory; at least one memory controller; and at least one input and output (I / O) controller.

7. The system-on-chip according to claim 1, wherein for the test step, the in-system component tester is further configured to instruct the one or more wrappers to isolate the connection of the at least one functional component selected as the component under test from the system bus.

8. The system-on-chip according to claim 1, wherein the in-system component tester includes at least one of the following: a detector configured to monitor the states of the plurality of functional components; a scheduler configured to manage the operations of the in-system component tester; a generator configured to generate test input data; and A tester configured to import the test input data into the component under test and analyze the test results obtained from the component under test that processes the test input data.

9. The system-on-chip according to claim 8, wherein, the test input data is predefined test data or a random bit stream generated based on a seed.

10. The system-on-chip according to claim 1, wherein the on-chip component tester is further configured to: After the test step is completed, analyze the test results obtained from at least one functional component selected as the component under test; and Based on the at least one functional component being analyzed as normal, allow the at least one functional component to be connected to the system bus or another system connection.

11. The system-on-chip according to claim 1, wherein the on-chip component tester is further configured to deactivate the at least one functional component based on the at least one functional component being analyzed as defective.

12. The system-on-chip according to claim 11, which further includes: A field programmable gate array (FPGA) configured to mimic the at least one functional component analyzed as defective.

13. The system-on-chip according to claim 12, wherein the field programmable gate array has an address that is revoked and replaced by the address of the at least one functional component analyzed as defective.

14. The system-on-chip according to claim 11, wherein the deactivation step includes revoking the address of the at least one functional component analyzed as defective.

15. The system-on-chip according to claim 11, wherein, the deactivation step includes powering off the at least one functional component analyzed as defective.

16. The system-on-chip according to claim 11, wherein, the deactivation step includes isolating the at least one functional component analyzed as defective from the system bus by cutting off the system bus connection to the at least one functional component analyzed as defective.

17. The system-on-chip according to claim 11, wherein, the plurality of functional components includes spare components of the at least one functional component analyzed as defective; and wherein the on-chip component tester is further configured to enable the spare components.

18. The system-on-chip according to claim 1, wherein the test step is repeatedly performed before and after the system-on-chip is delivered from the factory, and verifies whether there are defects in the manufacturing of the system-on-chip, whether it has been damaged, or whether it has been broken.

19. The system-on-chip according to claim 1, wherein the test step includes a scan test different from a functional test, and wherein, for the scan test, the on-chip component tester is further configured to: Connect multiple flip-flops in each component under test to each other, Import a test input into at least one flip-flop, and Obtain test results from the operation of the combinational logic of the flip-flop to analyze whether the component under test is defective or normal during runtime.

20. The system-on-chip according to claim 1, wherein the plurality of functional components includes a neural processing unit (NPU), and the neural processing unit includes a plurality of arrays of processing elements and is configured to select and test at least one of the plurality of arrays of processing elements.

21. A system-on-chip (SoC) for testing components in a system during runtime, the system-on-chip comprising: a plurality of functional components that communicate with each other via a system bus; one or more wrappers, each of the one or more wrappers being connected to one of the plurality of functional components; and an in-system component tester (ICT) configured to: when at least one functional component is monitored to be in an idle state via the one or more wrappers, select the at least one functional component in the idle state as a component under test (CUT), and based on detecting an access to the selected at least one functional component, interrupt the test of the at least one functional component selected as the component under test and allow the selected at least one functional component to be connected to the system bus.

22. The system-on-chip according to claim 21, wherein the in-system component tester is further configured to, after allowing the at least one functional component to be connected to the system bus, if it is confirmed that the at least one functional component is again in the idle state, re-select the at least one functional component as the component under test.

23. A method for testing components in a system-on-chip (SoC) during runtime, the system-on-chip comprising: a plurality of functional components, each of the plurality of functional components including circuitry; a system bus configured to allow the plurality of functional components to communicate with each other; one or more wrappers, each of the one or more wrappers being connected to one of the plurality of functional components; and an in-system component tester (ICT) configured to execute the method, the method comprising: monitoring the states of the plurality of functional components via the one or more wrappers; selecting at least one of the plurality of functional components in an idle state as a component under test (CUT); testing the at least one functional component selected as the component under test via the one or more wrappers; based on detecting a conflict with an access from the system bus to the at least one functional component selected as the component under test, interrupting the test step of the at least one functional component selected as the component under test; and based on the interrupt step, allowing the at least one functional component to be connected to the system bus.

24. The method according to claim 23, which further comprises: after allowing the at least one functional component to be connected to the system bus, if it is confirmed as a result of the monitoring step that the at least one functional component is again in the idle state, returning to the selecting step.

25. The method according to claim 23, which further comprises: isolating the connection of the at least one functional component selected as the component under test from the system bus.

26. The method according to claim 23, further comprising: after the testing step is completed, analyzing test results obtained from the at least one functional component selected as the component to be tested; and based on the at least one functional component being analyzed as normal, allowing the at least one functional component to be connected to the system bus.

27. The method according to claim 23, further comprising: based on the at least one functional component being analyzed as defective, deactivating the at least one functional component.

Citation Information

Patent Citations

  • Self-test during idle cycles for shader core of GPU

    CN111417932A