Fault processing method and apparatus, device, and medium

By constructing a grid-like topology in the data processing device, dividing it into unit groups and isolating faulty units, the problem of wasted hardware resources in the prior art is solved, achieving more efficient utilization of hardware resources and ensuring data processing performance.

CN119645691BActive Publication Date: 2026-03-27KUNLUNXIN TECHNOLOGY (BEIJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-11
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

In the existing technology, the fault detection methods of data processing devices are coarse-grained and fail to make full use of non-faulty areas, resulting in a waste of hardware resources.

Method used

By constructing a grid-like topology in the data processing device, dividing it into unit groups, determining the validity of each unit group, isolating faulty units and placing them in a bypass state, retaining fault-free areas, configuring path selectors to achieve bypass connections, and using spare sub-units to replace faulty sub-units, hardware resource utilization is optimized.

Benefits of technology

While isolating faulty units, the non-faulty areas are fully utilized to avoid wasting hardware resources, ensure the processing performance of the data processing device, and improve the availability and robustness of the data processing device.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119645691B_ABST
    Figure CN119645691B_ABST
Patent Text Reader

Abstract

The present disclosure provides a fault processing method and device applied to a data processing apparatus, equipment and a medium, relates to the technical field of computers, and particularly relates to the fields of data processing, fault processing and chip technology. The implementation scheme is as follows: a plurality of unit groups are determined based on the grid-shaped topology of the data processing apparatus; it is determined whether each data processing unit is a faulty unit; for each unit group, in response to determining that the unit group does not include a faulty unit, the unit group is determined to be a valid group; in response to determining that the number of valid groups meets a preset condition, a configuration operation is sequentially performed on the plurality of unit groups, the configuration operation comprising: for each unit group, in response to determining that the unit group includes a faulty unit that is not placed in a bypass state, a plurality of data processing units in the unit group are placed in a bypass state, wherein the data input port and the data output port of the data processing unit placed in the bypass state are in short-circuit connection; and in response to determining that the number of valid groups does not meet the preset condition, the data processing apparatus is determined to be a faulty apparatus.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of computer, in particular to the technical field of data processing, fault processing and chip, and more particularly to a fault processing method and device applied to a data processing apparatus, an electronic device, a computer readable storage medium and a computer program product. BACKGROUND

[0002] Artificial intelligence is a discipline that studies enabling computers to simulate some thinking processes and intelligent behaviors of human beings (such as learning, reasoning, thinking, planning, etc.), which has both hardware and software technologies. Artificial intelligence hardware technology generally includes technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing, etc.; artificial intelligence software technology mainly includes computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, knowledge graph technology, etc.

[0003] With the development of artificial intelligence technology, more and more applications based on artificial intelligence technology have achieved far better results than traditional algorithms. Deep learning is a data-intensive algorithm and a computing-intensive algorithm, in order to improve the training speed and inference speed of large-scale deep learning models, it is necessary to make more full use of the hardware resources of the data processing apparatus and reduce the computing power cost.

[0004] The methods described in this section can not necessarily be the methods previously conceived or adopted. Unless otherwise indicated, nothing in this section should be assumed to be prior art merely because it is included in this section. Similarly, matters discussed in this section should not be assumed to have been admitted to be prior art to any application unless otherwise indicated. SUMMARY

[0005] The present disclosure provides a fault processing method and device applied to a data processing apparatus, an electronic device, a computer readable storage medium and a computer program product.

[0006] According to an aspect of the present disclosure, there is provided a fault processing method applied to a data processing device, the data processing device comprising a plurality of data processing units constituting a mesh-shaped topology, the mesh-shaped topology comprising a plurality of rows and a plurality of columns, each of the plurality of data processing units comprising a data input port, a data processing region and a data output port connected in sequence, the method comprising: determining a plurality of unit groups based on the mesh-shaped topology, wherein the plurality of data processing units included in each unit group is capable of constituting a row or a column in the mesh-shaped topology; determining whether each of the plurality of data processing units is a faulty unit; for each of the plurality of unit groups, in response to determining that the faulty unit is not included in the unit group, determining the unit group as a valid group; in response to determining that a number of the valid groups meets a preset condition, sequentially performing a configuration operation on the plurality of unit groups, the configuration operation comprising: for each unit group, in response to determining that the faulty unit not being in a bypass state is included in the unit group, setting the plurality of data processing units in the unit group to a bypass state, wherein the data input port and the data output port of the data processing unit set to the bypass state are in short-circuit connection; and in response to determining that the number of the valid groups does not meet the preset condition, determining that the data processing device is a faulty device.

[0007] According to an aspect of the present disclosure, there is provided a fault processing device applied to a data processing device, the data processing device comprising a plurality of data processing units constituting a mesh-shaped topology, the mesh-shaped topology comprising a plurality of rows and a plurality of columns, each of the plurality of data processing units comprising a data input port, a data processing region and a data output port connected in sequence, the device comprising: a first determining unit configured to determine a plurality of unit groups based on the mesh-shaped topology, wherein the plurality of data processing units included in each unit group is capable of constituting a row or a column in the mesh-shaped topology; a second determining unit configured to determine whether each of the plurality of data processing units is a faulty unit; a third determining unit configured to, for each of the plurality of unit groups, in response to determining that the faulty unit is not included in the unit group, determine the unit group as a valid group; a first configuration unit configured to, in response to determining that a number of the valid groups meets a preset condition, sequentially perform a configuration operation on the plurality of unit groups, the configuration operation comprising: for each unit group, in response to determining that the faulty unit not being in a bypass state is included in the unit group, setting the plurality of data processing units in the unit group to a bypass state, wherein the data input port and the data output port of the data processing unit set to the bypass state are in short-circuit connection; and a fourth determining unit configured to, in response to determining that the number of the valid groups does not meet the preset condition, determine that the data processing device is a faulty device.

[0008] According to an aspect of the present disclosure, an electronic device is provided, including at least one processor; and a memory connected with the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the above fault processing method.

[0009] According to an aspect of the present disclosure, a non-transitory computer readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to perform the above fault processing method.

[0010] According to an aspect of the present disclosure, a computer program product is provided, including a computer program, wherein the computer program, when executed by a processor, can implement the above fault processing method.

[0011] According to one or more embodiments of the present disclosure, hardware resources of a data processing apparatus can be more fully utilized.

[0012] It should be understood that the content described in this section is not intended to identify key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become apparent through the following description. BRIEF DESCRIPTION OF DRAWINGS

[0013] The accompanying drawings, which are included to provide a further understanding of the application and are incorporated in and constitute a part of this specification, illustrate embodiments of the application and together with the description serve to explain exemplary implementations of the application. The illustrated embodiments are merely examples and do not limit the scope of the claims. In all the drawings, like reference numerals refer to like elements throughout the accompanying drawings.

[0014] Figure 1 A flow chart of a fault processing method applied to a data processing apparatus according to an exemplary embodiment of the present disclosure is shown;

[0015] Figure 2 A structural schematic diagram of a data processing system according to an exemplary embodiment of the present disclosure is shown;

[0016] Figure 3 A structural schematic diagram of a data processing apparatus according to an exemplary embodiment of the present disclosure is shown;

[0017] Figure 4 A structural schematic diagram of a data processing unit according to an exemplary embodiment of the present disclosure is shown;

[0018] Figure 5 A structural schematic diagram of a data processing area according to an exemplary embodiment of the present disclosure is shown;

[0019] Figure 6A structural block diagram of a fault processing device applied to a data processing device according to an example embodiment of the present disclosure is shown.

[0020] Figure 7 A structural block diagram of an example electronic device that can be used to implement an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0021] Exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings, which include various details of the embodiments of the present disclosure to assist in understanding them. These should be considered in a descriptive sense only and not limiting. Therefore, it will be recognized by those of ordinary skill that various changes in form and details can be made to the embodiments described herein without departing from the scope of the present disclosure. As such, the scope of the present disclosure should not be limited to the embodiments described in the following description, but should be given the full scope contemplated at this time. In addition, it will be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. It will be understood that the use of the singular herein includes the plural unless the context clearly dictates otherwise. The use of the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0022] In the present disclosure, the terms "first", "second", and the like are used to describe various elements only and do not intend to limit the positional relationship, the time sequence relationship, or the importance relationship of these elements. Such terms are only used to distinguish one element from another element. In some examples, the first element and the second element can refer to the same instance of the element, and in some cases, based on the context of the description, they can also refer to different instances.

[0023] The terms used in the description of various described examples in the present disclosure are only for the purpose of describing specific examples and are not intended to be limiting. Unless the context clearly dictates otherwise, if the number of elements is not specifically limited, the element can be one or more. In addition, the term "and / or" used in the present disclosure encompasses any one of the listed items and all possible combinations thereof.

[0024] Generally, a large number of memory and computing units are integrated in a chip for implementing data processing functions. In order to improve the computing power of a single chip, a large number of identical memory units and computing units are integrated therein. In the related art, a complete chip for implementing data processing functions (i.e., a data processing device) is usually detected, and when a fault is found, the data processing device is determined to be a faulty device. The granularity of this detection method is relatively coarse, and the non-faulty areas in the data processing device are not fully utilized.

[0025] Based on this, the disclosure provides a fault processing method applied to a data processing device, the data processing device comprising a plurality of computing units constituting a grid-shaped topology, the validity of each row and each column in the grid-shaped topology is determined, and when there is a faulty computing unit in a certain row or a certain column, the computing units in the row or the column are bypassed to obtain a fault-free data processing device meeting a preset condition. By applying the method, the computing power of the data processing device can be ensured on the premise of isolating the faulty data processing unit, the non-faulty area in the data processing device is fully utilized, and the waste of hardware resources is avoided.

[0026] Embodiments of the disclosure will be described in detail below with reference to the accompanying drawings.

[0027] Figure 1 A flowchart of a fault processing method 100 applied to a data processing device according to an exemplary embodiment of the disclosure is shown, the data processing device comprising a plurality of data processing units constituting a grid-shaped topology, the grid-shaped topology comprising a plurality of rows and a plurality of columns, each data processing unit in the plurality of data processing units comprising a data input port, a data processing region and a data output port connected in sequence. As shown, Figure 1 The method 100 comprises:

[0028] Step S101, determining a plurality of unit groups based on the grid-shaped topology, wherein the plurality of data processing units included in each unit group can constitute a row or a column in the grid-shaped topology;

[0029] Step S102, determining whether each data processing unit in the plurality of data processing units is a faulty unit;

[0030] Step S103, for each unit group in the plurality of unit groups, in response to determining that the faulty unit is not included in the unit group, determining that the unit group is a valid group;

[0031] Step S104, in response to determining that the number of valid groups meets a preset condition, sequentially performing a configuration operation for the plurality of unit groups, the configuration operation comprising: for each unit group, in response to determining that the faulty unit not in the bypass state is included in the unit group, placing the plurality of data processing units in the unit group in a bypass state, wherein the data input port and the data output port of the data processing unit placed in the bypass state are in short-circuit connection; and

[0032] Step S105, in response to determining that the number of valid groups does not meet the preset condition, determining that the data processing device is a faulty device.

[0033] By applying the fault processing method 100, the unit groups can be divided based on the grid topology formed by the plurality of data processing units of the data processing apparatus, the validity of each unit group is determined, and the faulty data processing unit can be isolated in units of unit groups, so as to retain the non-faulty data processing area meeting the condition, fully utilize the non-faulty area, ensure the processing performance of the data processing apparatus after isolation of the fault, and avoid waste of hardware resources.

[0034] In some examples, the data processing apparatus includes MxN computing units in a grid topology, based on which (M+N) unit groups corresponding to M rows and N columns in the grid topology can be determined. When there is a faulty data processing unit in a unit group, all the data processing units in the unit group are bypassed to obtain a non-faulty apparatus meeting the preset condition. In an example, the preset condition can be configured to indicate that the difference between the number of valid groups and the number of all unit groups should not be greater than a preset threshold value. For example, when the preset threshold value is 1, it can be ensured that the data processing apparatus after isolation of the fault can at least form a grid topology of (M-1)xN or a grid topology of Mx(N-1). Thus, after completing the fault processing, the grid topology can continue to support parallel data processing normally, and at the same time, the non-faulty area in the data processing apparatus is fully utilized, the available computing power is retained within a preset range, and waste of hardware resources is avoided.

[0035] In some examples, the order of performing the configuration operation on the plurality of unit groups can be set in advance according to requirements in step S104. It can be understood that each faulty unit in the grid topology corresponds to two unit groups. By performing the configuration operation on the unit groups in the preset order, redundant isolation operations can be avoided to retain more computing power as much as possible. In an example, the order can be set in advance to perform the configuration operation on the unit group with fewer data processing units first. For example, when a faulty unit corresponds to a unit group A including 5 data processing units and a unit group B including 3 data processing units, the bypass configuration operation can be performed on the unit group B first to bypass the faulty unit, so that the bypass configuration operation does not need to be performed on the unit group A to retain more computing power in the unit group A.

[0036] In some examples, the data processing apparatus can include a graphics processing unit (GPU), a central processing unit (CPU), or various types of computing chips, computing arrays, and other logic devices, such as a field programmable gate array (FPGA). The present disclosure does not limit the same.

[0037] According to some embodiments, the data processing unit further comprises a path selector, the data output port comprises a first sub-output port and a second sub-output port, wherein the first sub-output port is connectable to the data processing region through a first path of the path selector, the second sub-output port is short-circuit connectable to the data input port through a second path of the path selector, and wherein the bypassing the plurality of data processing units in the unit group in step S104 comprises configuring the state of the path selector of the plurality of data processing units in the unit group to select the second path. In this way, the first path capable of supporting normal data processing function and the second path for realizing the bypass function of fault isolation can be configured for the data processing unit by using the path selector, and the bypass state of the data processing unit can be configured more conveniently by using the path selector.

[0038] In some examples, the bypass state of the data processing unit can also be configured in other ways, for example, a short-circuit branch comprising a controlled switching device can be arranged between the data input port and the data output port of the data processing unit. The on-off of the short-circuit branch can be controlled by configuring the control signal of the controlled switching device, so as to realize the short-circuit connection between the data input port and the data output port.

[0039] According to some embodiments, the method 100 further comprises: storing the state information of the path selectors of the plurality of data processing units in the data processing device into a configuration table; and in response to receiving an initialization instruction for the data processing device, configuring the states of the path selectors of the plurality of data processing units based on the configuration table. In this way, after completing the fault processing of the data processing device, the states of the respective path selectors can be stored by using the configuration table, so that the respective path selectors can be directly configured based on the configuration table in the process of applying the device to realize data processing function, and the available data processing device can be obtained conveniently and efficiently.

[0040] In some examples, the data processing region in the data processing unit can be further divided to obtain a plurality of sub-units.

[0041] Based on this, in some embodiments, each data processing unit of the plurality of data processing units comprises a plurality of initial sub-units and at least one backup sub-unit, the backup sub-unit is not connected with the plurality of initial sub-units, and wherein whether each data processing unit of the plurality of data processing units is a faulty unit is determined by the following process: detecting whether each initial sub-unit of the plurality of initial sub-units and each backup sub-unit of the at least one backup sub-unit is a faulty sub-unit; in response to determining that the number of faulty sub-units in the plurality of initial sub-units is not more than the number of non-faulty sub-units in the at least one backup sub-unit, for each initial sub-unit of the plurality of initial sub-units, in response to determining that the initial sub-unit is a faulty sub-unit, replacing the faulty sub-unit with a non-faulty sub-unit of the at least one backup sub-unit; and in response to determining that the number of faulty sub-units in the plurality of initial sub-units is more than the number of non-faulty sub-units in the at least one backup sub-unit, determining that the data processing unit is a faulty unit. In this way, by configuring a first number of backup sub-units for a data processing unit comprising a plurality of initial sub-units, when a fault occurs in an initial sub-unit and the number of faulty initial sub-units is less than the number of available backup sub-units, the faulty sub-unit can be directly replaced with a backup sub-unit, the robustness of the data processing unit is improved through redundant backup, and the availability of the data processing device is further improved.

[0042] In some examples, each initial sub-unit in the data processing unit comprises a sequentially connected sub-unit input port, a sub-processing region, and a sub-unit output port, and the plurality of initial sub-units can be interconnected through the connection between the sub-unit input port and the sub-unit output port. Understandably, in the initial state without performing fault processing, the backup sub-unit is not interconnected with any initial sub-unit. By performing the above-mentioned fault processing steps, i.e., by reconfiguring the connection relationship between the input port sub-unit number input port, the sub-unit output port of the initial sub-unit, and the backup sub-unit, the replacement of the faulty sub-unit with the backup sub-unit can be achieved.

[0043] In some examples, each initial sub-unit can comprise a first sub-unit input port, a second sub-unit input port, a first sub-unit output port, and a second sub-unit output port, and a pass selector can be used to configure a third pass sequentially connecting the first sub-unit input port, the sub-processing region, and the first sub-unit output port, and a fourth pass sequentially connecting the second sub-unit input port, the backup sub-unit, and the second sub-unit output port. In this way, the replacement of the faulty sub-unit with the backup sub-unit can be conveniently and efficiently achieved by configuring the state of the pass selector.

[0044] In some examples, the state information of the path selectors corresponding to each of the plurality of initial sub-units in the data processing unit can be stored into a configuration table, so that the connection relationship between the initial sub-units and the standby sub-units can be directly configured based on the configuration table, and the available data processing unit can be conveniently and efficiently obtained.

[0045] In some examples, a connection branch including a controlled switching device can be arranged between the sub-unit input port and the sub-unit output port of the initial sub-unit and the standby sub-unit. By configuring the control signal of the controlled switching device, the connection branch can be controlled to be connected or disconnected, so as to implement the replacement step described above.

[0046] According to some embodiments, the step of detecting whether each of the plurality of initial sub-units and each of the at least one standby sub-unit is a faulty sub-unit includes: determining input data and reference output data corresponding to the input data; for each of the plurality of initial sub-units, performing data processing based on the input data by using the initial sub-unit to obtain test output data; and determining whether the initial sub-unit is a faulty sub-unit based on the test output data and the reference output data. In this way, the data processing task can be performed by using the sub-unit, and whether the sub-unit is faulty can be conveniently and efficiently determined by comparing the reference output data and the test output data.

[0047] In some examples, the input data can be transmitted into a data processing model corresponding to the hardware function of the data processing device established in the software environment. The data processing model can be a mathematical algorithm model or a chip circuit model that can be used for simulation. The model is used to simulate the function of the data processing device to obtain reference output data corresponding to the input data as a reference value of the test output data to be verified.

[0048] According to some embodiments, the step of determining that the data processing device is a faulty device in response to determining that the number of the valid groups does not meet the preset condition in the step S105 includes: determining a first proportion of the valid units in the plurality of data processing units that are not placed in the bypass state based on the number of the valid groups; and determining that the data processing device is a faulty device in response to determining that the first proportion is less than a preset proportion. In this way, the proportion of the available data processing units to the total data processing units can be limited by configuring the preset condition, so as to ensure that the computing power of the data processing device after fault processing is not lower than a preset threshold, and the normal execution of the data processing task is ensured.

[0049] Figure 2 A structural schematic diagram of a data processing system 200 according to an example embodiment of the present disclosure is shown. As shown in FIG. 1, the data processing system 200 includes a plurality of data processing units 210, a data processing device 220, and a data processing task management device 230. Figure 2As shown, the data processing system 200 can be considered as an integrated data processing chip, which includes multiple identical data processing devices 300, a storage unit 210 shared by the multiple data processing devices 300, and a communication unit 220. The multiple data processing devices 300 may be, for example, multiple identical computing cores.

[0050] In some examples, when testing an integrated data processing chip, the fault handling method 100 described above can be applied to each data processing device 300 separately to retain as many usable data processing devices 300 as possible. Once the fault handling process is completed, data processing chips whose number of usable data processing devices 300 meets a preset condition can be considered usable chips to improve chip yield.

[0051] Figure 3 A schematic diagram of the structure of a data processing apparatus 300 according to an exemplary embodiment of the present disclosure is shown. Figure 3 As shown, the data processing device includes 16 data processing units 400 capable of forming a 4×4 grid topology. By applying the above-described step S101, eight unit groups corresponding to four rows and four columns can be obtained. When a faulty data processing unit exists in a unit group, all data processing units in that unit group are bypassed to obtain a fault-free device that meets preset conditions. In one example, the preset condition can be that the number of valid groups is not less than 7, that is, the computing power of the retained data processing units is not less than 75% of the initial value. When a faulty data processing unit exists, by executing the fault handling method 100, a fault-free device capable of forming a 3×4 or 4×3 grid topology can be obtained.

[0052] Figure 4 A schematic diagram of the structure of a data processing unit 400 according to an exemplary embodiment of the present disclosure is shown. Figure 4 As shown, the data processing unit 400 includes a data input port 410, a data processing area 420, and a data output port 430 connected in sequence. The data output port 430 includes a first sub-output port 431 and a second sub-output port 432. By configuring a path selector for the data processing unit 400, a first path capable of supporting normal data processing functions and a second path for bypassing fault isolation can be configured for the data processing unit.

[0053] Figure 5A schematic diagram of the structure of a data processing region 420 according to an exemplary embodiment of the present disclosure is shown. As described above, by further dividing the data processing region, a plurality of initial sub-units 421 can be obtained. By applying the method described in the present disclosure, a certain number of spare sub-units 422 can be further configured for the data processing region. By using the spare sub-units 422 to replace the faulty initial sub-units 421, the robustness of the data processing unit is improved through redundancy backup.

[0054] In some examples, after completing the fault handling process for the data processing system 200, a configuration table can be used to store information about the available data processing devices 300, the status information of the path selectors of each data processing unit 400, and the status information of the path selectors of each initial subunit 421. This realizes the use of a configuration table to store the architecture of the data processing system after the fault is isolated, so as to conveniently and efficiently obtain the available data processing system.

[0055] According to one aspect of this disclosure, a fault handling apparatus for use in a data processing apparatus is also provided. Figure 6 A structural block diagram of a fault handling apparatus 600 applied to a data processing apparatus according to an exemplary embodiment of the present disclosure is shown. The data processing apparatus includes a plurality of data processing units constituting a grid-like topology, the grid-like topology including a plurality of rows and a plurality of columns. Each of the plurality of data processing units includes a data input port, a data processing area, and a data output port connected in sequence. Figure 6 As shown, the device 600 includes:

[0056] The first determining unit 601 is configured to determine multiple unit groups based on the grid topology, wherein the multiple data processing units included in each unit group can constitute a row or a column in the grid topology;

[0057] The second determining unit 602 is configured to determine whether each of the plurality of data processing units is a faulty unit;

[0058] The third determining unit 603 is configured to determine a unit group as a valid group in response to determining that the faulty unit is not included in the unit group for each of the plurality of unit groups;

[0059] The first configuration unit 604 is configured to, in response to determining that the number of valid groups meets the preset condition, sequentially perform a configuration operation on the plurality of unit groups, the configuration operation comprising: for each unit group, in response to determining that the unit group includes a faulty unit that is not in the bypass state, setting a plurality of data processing units in the unit group to the bypass state, wherein the data input port and the data output port of the data processing unit set to the bypass state are in short circuit connection; and

[0060] The fourth determination unit 605 is configured to, in response to determining that the number of valid groups does not meet the preset condition, determine that the data processing device is a faulty device.

[0061] According to some embodiments, the data processing unit further comprises a pass selector, the data output port comprises a first sub-output port and a second sub-output port, wherein the first sub-output port is connectable to the data processing region through a first pass of the pass selector, and the second sub-output port is in short circuit connection with the data input port through a second pass of the pass selector, and wherein the first configuration unit 604 is configured to configure the state of the pass selector of the plurality of data processing units in the unit group to select the second pass.

[0062] According to some embodiments, the device 600 further comprises a storage unit configured to store state information of the pass selector of the plurality of data processing units in the data processing device to a configuration table; and

[0063] The second configuration unit is configured to, in response to receiving an initialization instruction for the data processing device, configure the state of the pass selector of the plurality of data processing units based on the configuration table.

[0064] According to some embodiments, each data processing unit of the plurality of data processing units comprises a plurality of initial sub-units and at least one standby sub-unit, the standby sub-unit is not connected to the plurality of initial sub-units, and wherein the second determination unit 602 comprises: a detection sub-unit configured to detect whether each initial sub-unit of the plurality of initial sub-units and each standby sub-unit of the at least one standby sub-unit is a faulty sub-unit; a replacement sub-unit configured to, in response to determining that the number of faulty sub-units in the plurality of initial sub-units is not more than the number of non-faulty sub-units in the at least one standby sub-unit, for each initial sub-unit of the plurality of initial sub-units, in response to determining that the initial sub-unit is a faulty sub-unit, replace the faulty sub-unit with a non-faulty sub-unit of the at least one standby sub-unit; and a determination sub-unit configured to, in response to determining that the number of faulty sub-units in the plurality of initial sub-units is more than the number of non-faulty sub-units in the at least one standby sub-unit, determine that the data processing unit is a faulty unit.

[0065] According to some embodiments, the detecting subunit is configured to: determine input data and reference output data corresponding to the input data; for each initial subunit in the plurality of initial subunits, perform data processing based on the input data by the initial subunit to obtain test output data; and determine whether the initial subunit is a faulty subunit based on the test output data and the reference output data.

[0066] According to some embodiments, the fourth determining unit 605 is configured to: determine, based on the number of the valid groups, a first proportion of valid units in the plurality of data processing units that are not placed in a bypass state; and in response to determining that the first proportion is less than a preset proportion, determine that the data processing apparatus is a faulty apparatus.

[0067] It should be understood that, Figure 6 The operations of the various units of the fault detection apparatus 600 shown in FIG. 6 can correspond to the various steps in the fault detection method 100 described above. Figure 1 The operations, features and advantages described above for the method 100 also apply to the apparatus 600 and the various units included therein. For the sake of brevity, certain operations, features and advantages are not described again here.

[0068] According to an aspect of the present disclosure, there is also provided an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the fault detection method 100 described above.

[0069] According to an aspect of the present disclosure, there is also provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to perform the fault detection method 100 described above.

[0070] According to an aspect of the present disclosure, there is also provided a computer program product comprising a computer program, wherein the computer program, when executed by a processor, implements the fault detection method 100 described above.

[0071] Reference is made to Figure 7The present invention describes a structural block diagram of an electronic device 700 that can serve as a server or client of the present disclosure, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0072] like Figure 7 As shown, device 700 includes a computing unit 701, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 702 or a computer program loaded from storage unit 708 into random access memory (RAM) 703. RAM 703 may also store various programs and data required for the operation of device 700. The computing unit 701, ROM 702, and RAM 703 are interconnected via bus 704. Input / output (I / O) interface 705 is also connected to bus 704.

[0073] Multiple components in device 700 are connected to I / O interface 705, including: input unit 706, output unit 707, storage unit 708, and communication unit 709. Input unit 706 can be any type of device capable of inputting information to device 700. Input unit 706 can receive input numerical or character information and generate key signal inputs related to user settings and / or function control of the electronic device, and may include, but is not limited to, a mouse, keyboard, touchscreen, trackpad, trackball, joystick, microphone, and / or remote control. Output unit 707 can be any type of device capable of presenting information, and may include, but is not limited to, a monitor, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 708 may include, but is not limited to, a hard disk and an optical disk. Communication unit 709 allows device 700 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth™ devices, 802.11 devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.

[0074] The computing unit 701 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 701 performs various methods and processes described above, such as the fault detection method 100. For example, in some embodiments, the fault detection method 100 can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded onto the RAM 703 and executed by the computing unit 701, one or more steps of the fault detection method 100 described above can be performed. Alternatively, in other embodiments, the computing unit 701 can be configured to perform the fault detection method 100 by any other appropriate means, such as by means of firmware.

[0075] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a complex programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0076] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces a means for implementing the functions / acts specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, partially on a machine and partially on a remote machine or entirely on a remote machine or server.

[0077] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0078] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0079] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.

[0080] The computer system can include clients and servers. This relationship can be. The servers are typically remote from the users who use clients to interact with these servers. These clients and servers are often interconnected via communica tion networks. The relationship of client and server arises by interplay of programs in their respective computers and the concomitant con nection of the computers by a communication network. The servers can be cloud servers, servers of a distributed system, or servers incorporating blockchain.

[0081] It should be understood that the various forms of flow shown above can be used to reorder, add, or delete steps. For example, the steps recited in the present disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved, which are not limited herein.

[0082] Although the embodiments or examples of the present disclosure have been described with reference to the accompanying drawings, it should be understood that the above-described methods, systems, and devices are merely exemplary embodiments or examples, and the scope of the present disclosure is not limited by these embodiments or examples. Various elements in the embodiments or examples can be omitted or replaced by equivalent elements thereof. In addition, each step can be performed in an order different from that described in the present disclosure. Further, various elements in the embodiments or examples can be combined in various ways. It is important that many of the elements described herein can be replaced by equivalent elements that appear after the present disclosure as technology evolves.

Claims

1. A fault handling method applied to a data processing device, the data processing device comprising a plurality of data processing units constituting a grid topology, the grid topology comprising a plurality of rows and a plurality of columns, each of the plurality of data processing units comprising a data input port, a data processing area, and a data output port connected in sequence, the method comprising: Based on the grid topology, multiple unit groups are determined, wherein the multiple data processing units included in each unit group can constitute a row or a column in the grid topology; Determine whether each of the plurality of data processing units is a faulty unit; For each of the plurality of unit groups, in response to determining that the faulty unit is not included in the unit group, the unit group is determined to be a valid group; In response to determining that the number of valid groups meets a preset condition, a configuration operation is sequentially performed on the plurality of unit groups. The configuration operation includes: for each unit group, in response to determining that the unit group includes a faulty unit that is not set to bypass mode, setting a plurality of data processing units in the unit group to bypass mode, wherein the data input port and data output port of the data processing unit set to bypass mode are short-circuited; and In response to determining that the number of valid groups does not meet the preset condition, the data processing device is determined to be a faulty device. Each of the plurality of data processing units includes a plurality of initial sub-units and at least one spare sub-unit, wherein the spare sub-unit is not connected to the plurality of initial sub-units, and wherein whether each data processing unit is a faulty unit is determined by the following process: Detect whether each of the plurality of initial sub-units and each of the at least one spare sub-unit is a faulty sub-unit; In response to determining that the faulty subunit among the plurality of initial subunits is no more than the non-faulty subunit among the at least one spare subunit, for each of the plurality of initial subunits, in response to determining that the initial subunit is a faulty subunit, the faulty subunit is replaced by a non-faulty subunit among the at least one spare subunit; and In response to determining that the number of faulty subunits in the plurality of initial subunits is greater than the number of non-faulty subunits in the at least one spare subunit, the data processing unit is determined to be a faulty unit. Furthermore, the step of sequentially performing configuration operations on the plurality of unit groups includes: when the number of processing units corresponding to the rows and columns of the grid topology is different, first performing the configuration operation on the unit group with the smaller number of corresponding processing units in the rows and columns.

2. The method as described in claim 1, wherein, The data processing unit further includes a path selector, and the data output port includes a first sub-output port and a second sub-output port. The first sub-output port can be connected to the data processing area through a first path of the path selector, and the second output port can be short-circuited to the data input port through a second path of the path selector. Furthermore, setting the plurality of data processing units in the unit group to a bypass state includes: Configure the path selector of multiple data processing units in this unit group to select the second path.

3. The method of claim 2, further comprising: The status information of the path selectors of multiple data processing units in the data processing device is stored in the configuration table; as well as In response to receiving an initialization command for the data processing device, the state of the path selectors of the plurality of data processing units is configured based on the configuration table.

4. The method of claim 1, wherein, The step of detecting whether each of the plurality of initial subunits and each of the at least one backup subunit is a faulty subunit includes: Determine the input data and the corresponding reference output data; For each of the plurality of initial sub-units, The initial subunit is used to perform data processing based on the input data to obtain test output data; and Based on the test output data and the reference output data, determine whether the initial subunit is a faulty subunit.

5. The method according to any one of claims 1-4, wherein, The step of determining that the data processing device is a faulty device in response to determining that the number of valid groups does not meet the preset condition includes: Based on the number of effective groups, a first proportion of the effective units that are not set to bypass state among the plurality of data processing units is determined; and In response to determining that the first ratio is less than a preset ratio, the data processing device is determined to be a faulty device.

6. A fault handling device for a data processing apparatus, the data processing apparatus comprising a plurality of data processing units constituting a grid topology, the grid topology comprising a plurality of rows and a plurality of columns, each of the plurality of data processing units comprising a data input port, a data processing area, and a data output port connected in sequence, the device comprising: The first determining unit is configured to determine multiple unit groups based on the grid topology, wherein the multiple data processing units included in each unit group can constitute a row or a column in the grid topology; The second determining unit is configured to determine whether each of the plurality of data processing units is a faulty unit; The third determining unit is configured to, for each of the plurality of unit groups, determine that the unit group is a valid group in response to determining that the faulty unit is not included in the unit group. A first configuration unit is configured to, in response to determining that the number of valid groups meets a preset condition, sequentially perform configuration operations on the plurality of unit groups. The configuration operations include: for each unit group, in response to determining that the unit group includes a faulty unit that is not set to a bypass state, setting a plurality of data processing units in the unit group to a bypass state, wherein the data input port and data output port of the data processing unit set to the bypass state are short-circuited; and The fourth determining unit is configured to determine the data processing device as a faulty device in response to determining that the number of valid groups does not meet the preset condition. Each of the plurality of data processing units includes a plurality of initial sub-units and at least one spare sub-unit, wherein the spare sub-unit is not connected to the plurality of initial sub-units, and wherein the second determining unit includes: The detection subunit is configured to detect whether each of the plurality of initial subunits and each of the at least one backup subunit is a faulty subunit. The replacement subunit is configured to, in response to determining that the faulty subunit among the plurality of initial subunits is no more than the non-faulty subunit among the at least one spare subunit, replace the faulty subunit with a non-faulty subunit among the at least one spare subunit for each of the plurality of initial subunits in response to determining that the initial subunit is a faulty subunit; and The subunit is configured to determine the data processing unit as a faulty unit in response to determining that more faulty subunits among the plurality of initial subunits are than non-faulty subunits among the at least one spare subunit. Furthermore, the first configuration unit is configured to, when the number of processing units corresponding to rows and columns of the mesh topology is different, first perform the configuration operation for the unit group for the row and column with fewer corresponding processing units.

7. The apparatus of claim 6, wherein, The data processing unit further includes a path selector, and the data output port includes a first sub-output port and a second sub-output port. The first sub-output port can be connected to the data processing area via a first path of the path selector, and the second sub-output port can be short-circuited to the data input port via a second path of the path selector. The first configuration unit is configured to: Configure the path selector of multiple data processing units in this unit group to select the second path.

8. The apparatus of claim 7, further comprising: The storage unit is configured to store the status information of the path selectors of a plurality of data processing units in the data processing device into a configuration table; as well as The second configuration unit is configured to, in response to receiving an initialization instruction for the data processing device, configure the state of the path selectors of the plurality of data processing units based on the configuration table.

9. The apparatus of claim 6, wherein, The detection subunit is configured as follows: Determine the input data and the corresponding reference output data; For each of the plurality of initial sub-units, The initial subunit is used to perform data processing based on the input data to obtain test output data; as well as Based on the test output data and the reference output data, determine whether the initial subunit is a faulty subunit.

10. The apparatus according to any one of claims 6-9, wherein, The fourth determining unit is configured as follows: Based on the number of effective groups, a first proportion of effective units that are not set to bypass state among the plurality of data processing units is determined. as well as In response to determining that the first ratio is less than a preset ratio, the data processing device is determined to be a faulty device.

11. An electronic device, comprising: At least one processor; as well as A memory that is communicatively connected to the at least one processor; in The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-5.

12. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-5.

13. A computer program product comprising a computer program, wherein, The computer program, when executed by a processor, implements the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • A fault tolerance method and system chip of an artificial intelligence module

    CN109902836A

  • MMC optimal redundancy configuration method based on NSGA-II

    CN112307618A