Computer system, its control method and arithmetic device
The computer system with trace data generation and timestamp values addresses the challenge of identifying failures and monitoring internal states in systems with multiple processing elements, enhancing system management and recovery efficiency.
Patent Information
- Application Number
- JP2023559357
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-11-12
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2041-11-12
AI Technical Summary
In computer systems with multiple processing elements, identifying failures within processing elements and grasping the internal state of the system is difficult due to independent data transfer without host unit involvement.
A computer system with a plurality of calculation units and a control unit that generates trace data including timestamp values and location information upon event detection, allowing for failure identification and internal state monitoring.
Enables easy identification of failure locations and internal data status, facilitating efficient system management and recovery.
Smart Images

Figure 0007790443000001 
Figure 0007790443000002 
Figure 0007790443000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a computer system having a plurality of processing units and a control method thereof. [Background technology]
[0002] Technological innovation is progressing in many fields, including machine learning, artificial intelligence (AI), and the Internet of Things (IoT), and by utilizing various information and data, services are being actively improved and value-added is being provided. Such processing requires a large amount of calculations, and therefore an information processing infrastructure is essential.
[0003] For example, Non-Patent Document 1 points out that attempts are being made to update existing information processing infrastructure, but current computers are unable to keep up with the rapidly increasing amount of data, and that in order to make progress in the future, "post-Moore technology" that goes beyond Moore's Law must be established.
[0004] As a post-Moore technology, for example, Non-Patent Document 2 discloses a technology called flow-centric computing. Flow-centric computing introduces a new concept of moving data to a location where a computing function is present and processing the data there, instead of the conventional computing concept of performing processing where the data is located.
[0005] To realize the above-mentioned flow-centric computing, not only is a broadband communication network necessary for data movement required, but it is also necessary to efficiently control computational resources in order to obtain the desired computing performance.
[0006] Flow-centric computing (for example, Non-Patent Document 2) discloses a method for linking multiple computing functions. [Prior art documents] [Non-patent literature]
[0007] [Non-Patent Document 1] “NTT Technology Report for Smart World 2020,” Nippon Telegraph and Telephone Corporation, 2020. https: / / www.rd.ntt / _assets / pdf / techreport / NTT_TRFSW_2020_EN_W.pdf. [Non-patent document 2] R. Takano and T. Kudoh, “Flow-centric computing leveraged by photonic circuit switching for the post-moore era,” Tenth IEEE / ACM International Symposium on Networks-on-Chip (NOCS), Nara, 2016, pp. 1-3. https: / / ieeexplore.ieee.org / abstract / document / 7579339. Summary of the Invention [Problem to be solved by the invention]
[0008] However, in a computer system in which multiple processing elements work together, it is difficult to identify a failure that occurs within a processing element because the processing elements independently transfer data between each other without going through the host unit.
[0009] Furthermore, it is difficult to grasp the internal state of a computer system, such as identifying the processing unit through which input data has passed at a certain time. [Means for solving the problem]
[0010] In order to solve the above-mentioned problems, the computer system according to the present invention includes a plurality of The computer includes a computing unit, and data is transferred between the computing units. The computing unit When an event is detected, a trace containing a timestamp value that is the detection time of the event is generated. Equipped with a trace section that records data The trace data includes information specifying a location where the event is detected. It is characterized by the following. The computer system according to the present invention is a computer system for processing input data. a system including a plurality of calculation units and a control unit connected to the plurality of calculation units and controlling the plurality of calculation units; a host unit for transferring the processed data between the plurality of calculation units, The calculation unit records the trace data when a predetermined event is detected from the input data. a trace unit for recording the trace data, the trace data including a timestamp value indicating the detection time of the event; and information specifying a location where the event is to be detected. The present invention is characterized by having the following.
[0011] A control method for a computer system according to the present invention is a control method for a computer system comprising a plurality of processing units and a host unit, the plurality of processing units each comprising an event generator, a timestamp unit, and a trace buffer, and data being input to the processing units, the control method comprising the steps of: the timestamp unit counting time based on an operating frequency of the processing units; the event generator detecting a predetermined event from the data, and generating a timestamp value in response to the detection of the event. and information specifying a location where the event is to be detected. and acquiring the Furthermore, an arithmetic device according to the present invention is an arithmetic device in a computer system, which is connected to a host unit and controlled by the host unit, and which includes a trace unit that receives data from another arithmetic device and records trace data including a timestamp value that is the detection time of a predetermined event upon detection of the event. The trace data includes information specifying a location where the event is detected. It is characterized by the following. [Effects of the Invention]
[0012] According to the present invention, it is possible to provide a computer system and a control method thereof that, when a failure occurs, can easily identify the location of the failure and grasp the internal data status at the time of the failure. [Brief explanation of the drawings]
[0013] [Figure 1] FIG. 1 is a block diagram showing the configuration of a computer system according to a first embodiment of the present invention. [Figure 2] FIG. 2 is a block diagram showing the configuration of the arithmetic unit in the computer system according to the first embodiment of the present invention. [Figure 3A] FIG. 3A is a diagram for explaining expansion and contraction of the processing units in the computer system according to the first embodiment of the present invention. [Figure 3B] FIG. 3B is a diagram for explaining expansion and contraction of the processing units in the computer system according to the first embodiment of the present invention. [Figure 4A] FIG. 4A is a flowchart illustrating a control method for a computer system according to a first embodiment of the present invention. [Figure 4B] FIG. 4B is a flowchart illustrating a control method for a computer system according to the first embodiment of the present invention. [Figure 4C] FIG. 4C is a diagram for explaining a control method of a computer system according to the first embodiment of the present invention. [Figure 5] FIG. 5 is a flowchart illustrating a control method for a computer system according to a second embodiment of the present invention. [Figure 6] FIG. 6 is a flowchart illustrating a control method for a computer system according to a third embodiment of the present invention. [Figure 7A] FIG. 7A is a diagram for explaining an example of a control method for the computer system according to the first embodiment of the present invention. [Figure 7B] FIG. 7B is a diagram for explaining an example of a control method for the computer system according to the first embodiment of the present invention. [Figure 8] FIG. 8 is a flowchart illustrating a control method for a computer system according to the second embodiment of the present invention. [Figure 9]FIG. 9 is a flowchart illustrating a control method for a computer system according to the second embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0014] First Embodiment A computer system and a control method thereof according to a first embodiment of the present invention will be described with reference to FIGS.
[0015] <Computer system configuration> As shown in FIG. 1, a computer system 10 according to this embodiment includes N calculation units 11_1 to 11_N (N is an integer equal to or greater than 1), an internal communication unit 13 connecting the calculation units 11_1 to 11_N, and a host unit 12 that sets and manages operating parameters for the calculation units 11_1 to 11_N.
[0016] The calculation units 11_1 to 11_N are configured by processors, accelerators, etc., and include trace units 14_1 to 14_N.
[0017] When the calculation units 11_1 to 11_N operate in conjunction with each other, the trace units 14_1 to 14_N record the event detection time based on the operating frequency of each calculation unit 11_1 to 11_N upon detection of a predetermined event at any observation point of each calculation unit 11_1 to 11_N.
[0018] Here, the trace units 14_1 to 14_N can record the event detection time for each data type or event type, and may further record any data.
[0019] In the interlocking of the calculation units 11_1 to 11_N, data processed by the calculation unit 11_1 is transferred to the calculation unit 11_2 via the internal communication unit 13. Subsequently, the data transfer is repeated and the data is transferred to the calculation unit 11_N.
[0020] The methods of linking the calculation units 11_1 to 11_N include a processing method of connecting multiple calculation units in series, a processing method of connecting multiple calculation units in parallel, a processing method that combines both, etc. By linking multiple calculation units, desired services are provided and applications are processed.
[0021] The calculation units 11_1 to 11_N (N is an integer equal to or greater than 1) have a function of executing predetermined calculation processing on input data input from outside the computer system 10. The calculation processing is, for example, general calculation processing such as processing, aggregation, and combination of input data, such as reducing or enlarging the image size when image data is input, detecting a specific object from image data, and decrypting or encrypting image data.
[0022] Furthermore, the arithmetic units 11_1 to 11_N may be added or removed regardless of whether the system is stopped or running. For example, this can be realized by using an FPGA, which is a device that allows dynamic reconfiguration of only a portion of the arithmetic units. Furthermore, as a method of implementing the arithmetic units 11_1 to 11_N, an accelerator card equipped with a dedicated circuit specialized for a specific operation may be added. Furthermore, a arithmetic unit may be provided with multiple arithmetic units that provide arithmetic functions.
[0023] The host unit 12 has a function of setting and managing operation parameters for the calculation units 11_1 to 11_N, and more specifically, a function of controlling the calculation units 11_1 to 11_N and a function of storing data. The operation parameters are, for example, information for identifying an algorithm when switching between multiple algorithms in image processing, such as coefficients and thresholds in the calculation process.
[0024] Furthermore, if it is possible to add or remove a processing unit even after the system has started operating, the host unit 12 manages the entire computer system 10, such as by setting circuit information for the processing unit to execute the desired processing content.
[0025] The internal communication unit 13 connects the arithmetic units 11_1 to 11_N and has a communication function for transmitting and receiving data between the arithmetic units 11_1 to 11_N. Specifically, examples of the internal communication unit 13 include a common communication standard such as PCIe or Ethernet and a physical configuration that satisfies the communication standard, that is, a PCIe switch or an Ethernet switch.
[0026] Furthermore, when the calculation units 11_1 to 11_N are provided with a plurality of calculation units that provide calculation functions, a plurality of the observation points may be provided in the calculation units 11_1 to 11_N.
[0027] <Configuration of the calculation unit> 2, the calculation unit 11_1 in the computer system 10 includes a plurality of (N) calculation units 15_1(1) to 15_N(1) and a tracing unit 14_1. Here, the number of calculation units may be one.
[0028] The trace unit 14_1 includes event generators 16_1_1 to 16_2_N, a time stamp unit 17, and a trace buffer 18.
[0029] The event generators 16_1_1 to 16_2_N are connected to the input and output sides of the arithmetic units 15_1(1) to 15_N(1), respectively. The outputs of the time stamp unit 17 are connected to the event generators 16_1_1 to 16_2_N. The outputs of the event generators 16_1_1 to 16_2_N are connected to the trace buffer 18.
[0030] Here, the event generators 16_1_1 to 16_2_N may be arranged on either the input side or the output side of the arithmetic units 15_1(1) to 15_N(1), and may be arranged at any location, and at least one unit may be arranged. Also, a plurality of trace buffers 18 may be arranged, and at least one unit may be arranged.
[0031] The event generators 16_1_1 to 16_2_N are inserted at any position in the calculation units 11_1 to 11_N, detect events (beginning and end of stream) for each type of data (user ID, session ID, stream ID, service ID), and generate an opportunity to record trace data including the detection time (hereinafter referred to as "timestamp value") in the trace buffer 18 described later.
[0032] Furthermore, the type of data is not limited to the above, and any information that can be used to organize data, such as header information of packets used to organize data or information contained in signals that run parallel to the data, can be applied.
[0033] The time stamp unit 17 includes at least one clock counter, synchronizes a plurality of event generators 16_1_1 to 16_2_N (observation points), and acquires time with the accuracy of the operating frequency of the arithmetic units 11_1 to 11_N. Here, the operating frequency (clock frequency) of the arithmetic units 11_1 to 11_N is usually about several nanoseconds when the function is realized using an FPGA (field-programmable gate array).
[0034] The trace buffer 18 records trace data when an event is detected by the event generators 16_1_1 to 16_2_N. Here, the trace data includes each detection time (timestamp value) acquired from the timestamp unit 17, an instance ID, an event type (event ID), a data type (TID), and arbitrary data. Here, the trace data only needs to have at least a timestamp value.
[0035] Furthermore, the trace buffer 18 provides a constant buffer capacity independent of the number of event generators.
[0036] Here, the timestamp value is a standardized value within the calculation unit (FPGA).
[0037] The instance ID is an ID that distinguishes the event generator instance and indicates the location (observation point) where the event is detected.
[0038] The event type (event ID) is an ID that distinguishes the event content. For example, it can be distinguished by the passage of the beginning or end of a stream. In addition, a flag or the like for event detection can be prepared at any point in the data, and the passage of that flag can be detected.
[0039] The arbitrary data is data that is normally processed by a computer system, such as image data, numerical data, and text data.
[0040] The data type is used to identify and classify the attributes of input data, for example, and is information attached to the data itself, such as a user ID, session ID, stream ID, or service ID. Information for identifying the data type does not necessarily have to be added to the packet header; for example, it may be uniquely defined in the packet payload. When a signal running parallel to the data is used inside the calculation unit, the parallel signal may be used to acquire the data type.
[0041] <Computer system operation> The operation of the computer system 10 according to this embodiment will be described below.
[0042] In the arithmetic unit 11_1 of the computer system 10, data is input to the arithmetic units 15_1(1) to 15_N(1). The input data is made up of various elements, and includes an event type (event ID), a data type (TID), and arbitrary data.
[0043] Here, data processed by the arithmetic unit 15_1(1) of the arithmetic unit 11_1 is transmitted and input to the arithmetic unit 15_1(2) of the arithmetic unit 11_2.
[0044] In the event generators 16_1_1 to 16_2_N, first, control signals (operations) of data input to the arithmetic units 11_1 to 11_N are observed.
[0045] When the event generators 16_1_1 to 16_2_N detect an event, the event generators 16_1_1 to 16_2_N acquire the event type, the data type, and any data from the input data. Here, the event occurrence is, for example, when the head of a stream passes, or when the head of a stream passes.
[0046] The event generators 16_1_1 to 16_2_N add the instance ID and the timestamp value transmitted from the timestamp unit 17 to the acquired event type, data type, and arbitrary data. As a result, the trace data is composed of the timestamp value, the instance ID, the event type, the data type, and arbitrary data.
[0047] Here, the trace data only needs to include at least a timestamp value, and based on the timestamp value, information within the computer system such as processing time, data flow rate, etc. can be ascertained. It is also possible to ascertain the time when a malfunction occurred.
[0048] Furthermore, since the trace data has an instance ID, it is possible to identify the location where the problem occurred.
[0049] Furthermore, the trace data includes an event type, so that the time when an event occurs can be known.
[0050] Furthermore, since the trace data has a data type, the operating status of each data type can be grasped, and this can be used to determine when to erase data (described later).
[0051] Furthermore, the trace data can contain arbitrary data, which can be used when restarting processing (described later).
[0052] Furthermore, since the trace data includes service priority information, it can be used to manage the trace data based on priority.
[0053] Finally, the event generators 16_1_1 to 16_2_N transmit the trace data to the trace buffer 18.
[0054] Furthermore, the time stamp unit 17 receives (writes) the count start or stop setting transmitted from the host unit 12.
[0055] This triggers the clock counter to start or stop counting.
[0056] The counting is started by the count start setting, and is executed based on the operating frequency of each of the arithmetic units 11_1 to 11_N, whereas the counting is stopped by the count stop setting.
[0057] When the event generators 16_1_1 to 16_2_N detect an event, they read out the time counted by the clock counter as a timestamp value, and transmit the timestamp value to the event generators 16_1_1 to 16_2_N.
[0058] In addition, events can be detected, for example, by determining whether a signal running parallel to the data is valid or not by checking whether it is ON / OFF, or by preparing a field for event detection in a specific area of the data and using the bit string of that field.
[0059] Here, the arithmetic units (FPGAs) are synchronized as necessary. As a method for synchronizing the arithmetic units, a signal for synchronization or a reset signal is input from the host unit to the arithmetic units to be synchronized.
[0060] Furthermore, when the host unit 12 reads out trace data from the trace buffer 18, it sends a reset signal to the time stamp unit 17 to reset the value of the clock counter.
[0061] In the trace buffer 18, the trace data received from the event generators 16_1_1 to 16_2_N is recorded and accumulated in the trace buffer 18.
[0062] Here, trace data transmitted from a plurality of event generators 16_1_1 to 16_2_N is recorded. Furthermore, writing and reading of trace data is performed in a common FIFO (First-In First-Out) for all TIDs.
[0063] Next, the host unit 12 reads the trace data from the trace buffer 18 .
[0064] Finally, the host unit 12 performs post-processing on the read (collected) data. Specifically, it searches (GREP) for each TID. Then, it is sorted by the timestamp unit 17 and visualized.
[0065] Here, an example has been shown in which the event generator acquires the event type, data type, and any data from the input data when an event is detected, but as described above, the computer system 10 can operate even if the event type, data type, and any data are not acquired, as long as at least a timestamp value is acquired.
[0066] <Addition and deletion of calculation units and calculation units> In the computer system 10, the arithmetic units 11_1 to 11_N can be added or deleted regardless of whether the computer system 10 is stopped or running.
[0067] Also, for example, in the calculation unit 11_1 shown in FIG. 3A, by newly arranging (adding) event generators 16_1_2 to 16_2_N and connecting them to the time stamp unit 17 and trace buffer 18, as shown in FIG. 3B, calculation units 15_1(1) to 15_N(1) can be added.
[0068] Moreover, in the arithmetic unit 11_1, by deleting the event generators 16_1_2 to 16_2_N, the arithmetic units 15_1(1) to 15_N(1) can be deleted.
[0069] When the computing units 15_1(1) to 15_N(1) are added, a plurality of event generators 16_1_1 to 16_2_N can be arranged at any positions in the tracing units 14_1 to 14_N.
[0070] Here, the event generators 16_1_1 to 16_2_N may be arranged both before the input and after the output of the arithmetic units 15_1(1) to 15_N(1), or may be arranged either before the input or after the output.
[0071] When the computing units 15_1(1) to 15_N(1) are deleted, the event generators 16_1_1 to 16_2_N before and after the computing units 15_1(1) to 15_N(1) in the tracing units 14_1 to 14_N may be deleted.
[0072] Here, the event generators at both the front stage of the input and the rear stage of the output of the computing units 15_1(1) to 15_N(1) to be deleted may be deleted, or either the front stage of the input or the rear stage of the output may be deleted.
[0073] According to this embodiment, multiple event generators can be added at any location. Therefore, by simply adding a new event generator, a new computing unit can be easily added inside the computing unit, making it possible to measure the processing time in the computing unit and collect trace data.
[0074] Furthermore, since multiple event generators can be deleted from any location, by deleting only the event generators, it is possible to easily delete a computing unit within the computing unit.
[0075] Furthermore, when adding a computing unit, the circuit size can be reduced and power consumption can be suppressed compared to when an event generator, a time stamp unit, and a trace buffer are all added.
[0076] Furthermore, when deleting a computing unit from the computing section, event generators located before and after the computing unit may also be deleted. At this time, trace data related to the event generators may also be deleted. Furthermore, trace data associated with events detected by the deleted event generators does not necessarily need to be deleted.
[0077] <First Example> A control method for the computer system 10 according to the first embodiment of the present invention will be described with reference to FIGS. 4A and 4B.
[0078] In this embodiment, in the computer system 10, trace data recorded by the trace units 14_1 to 14_N of the calculation units 11_1 to 11_N is used to efficiently resume processing when a malfunction occurs. Here, the malfunction may be a packet loss that may normally occur, a process stuck in a functional block inside the calculation units 11_1 to 11_N, or the like.
[0079] In the computer system 10, for example, the trace unit 14_1 of the calculation unit 11_1 records a timestamp value, an instance ID (information indicating a location where an event is detected), and arbitrary data as arbitrary trace data in the trace buffer 18.
[0080] As an example of a control method for the computer system 10 according to this embodiment, a case where the host unit 12 controls (manages) will be described with reference to FIG. 4A.
[0081] First, the host unit 12 monitors the trace buffer 18 of the computer system at a predetermined cycle and measures the processing time within the processing unit (step S11A).
[0082] Here, in measuring the processing time, first, the host unit 12 reads (acquires) from the trace buffer 18 the first timestamp value and the second timestamp value triggered by the detection of a predetermined event (first event, second event) by the event generator 16_1_1 (first event generator) on the input side of the arithmetic unit 15_1(1) and the event generator 16_2_1 (second event generator) on the output side of the arithmetic unit 15_1(1).
[0083] Subsequently, the host unit 12 calculates the processing time from the difference between the first time stamp value and the second time stamp value.
[0084] In this way, the processing time of any location (section, for example, arithmetic unit 15_1(1)) is obtained by the difference between the times (the first timestamp value and the second timestamp value) at which input data passes through each of the event generators placed at any location (section, for example, before and after arithmetic unit 15_1(1)).
[0085] Next, the measured processing time is compared with a predetermined threshold value (step S12A). If the processing time is found to be longer than the predetermined threshold value, it is determined that a malfunction has occurred.
[0086] Finally, if a problem occurs, the process detection point is identified from the instance ID, and processing is resumed using any data recorded in any of the trace buffers 18 before this process detection point (step S13A).
[0087] As an example of a control method for the computer system 10 according to this embodiment, a case where the arithmetic units 11_1 to 11_N perform control (management) will be described with reference to FIGS. 4B and 4C.
[0088] First, the calculation units 11_1 to 11_N record trace data and monitor the trace buffer 18 at a predetermined period, and measure the time (hereinafter referred to as the "processing time between calculation units") from when a timestamp value is recorded in the trace buffer 18 of any calculation unit (e.g., calculation unit 11_1) to when the data is input to the calculator 15_1(2) of the next-stage calculation unit (e.g., calculation unit 11_2) (step S11B).
[0089] In measuring the processing time between the calculation units, for example, as shown in FIG. 4C, first, the event generator 16_2_1 in the preceding stage of the trace buffer 18 of the calculation unit 11_1 acquires a timestamp value (first stamp value), which is then recorded in the trace buffer 18 of the calculation unit 11_1.
[0090] Subsequently, a notification signal is sent from the trace buffer 18 of the calculation unit 11_1 to the calculator 15_1(2) of the calculation unit 11_2 (dotted arrow in the figure), and this signal triggers data transfer from the calculator 15_1(1) of the calculation unit 11_1 to the calculator 15_1(2) of the calculation unit 11_2.
[0091] Subsequently, upon detection of an event in the data input to the calculator 15_1(2) of the calculation unit 11_2, the event generator 16_1_1 in the preceding stage of the calculator 15_1(2) of the calculation unit 11_2 acquires a timestamp value (second stamp value), which is recorded in the trace buffer 18 of the calculation unit 11_2.
[0092] The processing time between the calculation units is measured from the difference between the first stamp value and the second stamp value.
[0093] Next, the measured processing time is compared with a predetermined threshold value (step S12B), and if the measured time is longer than the predetermined threshold value, it is determined that a malfunction has occurred.
[0094] Finally, if a malfunction occurs, the process is resumed using any data recorded in any trace buffer 18 preceding the processing detection point identified by the instance ID recorded in the trace buffer 18 (step S13B).
[0095] Here, an example has been shown in which the processing time is measured using the timestamp value recorded in the trace buffer, but the processing time may also be measured directly using the timestamp value acquired by the event generator.
[0096] In this way, in the control method for a computer system according to this embodiment, an event generator placed at an arbitrary location detects a specified event, and in response to this, acquires one timestamp value, and an event generator placed at another location similarly acquires another timestamp value, and by calculating the difference between the one timestamp value and the other timestamp value, it is determined that a malfunction has occurred and processing is resumed.
[0097] According to the control method of the computer system of this embodiment, by using the trace data recorded by the trace units 14_1 to 14_N, when a malfunction occurs, it is possible to resume processing by tracing back to the point where it was operating normally, without restarting the processing from the beginning.
[0098] The trigger for restarting the processing of trace data is managed by the host unit 12. Alternatively, it may be managed by the calculation units 11_1 to 11_N.
[0099] Furthermore, the restart of processing does not necessarily have to be limited to a specific functional block, and for example, processing may be restarted by tracing back to the input of the calculation unit.
[0100] In this embodiment, an example has been shown in which the processing detection location is identified by the instance ID, but this is not limited to this. The processing detection location may also be derived using a preset processing speed of the arithmetic unit and a measured timestamp value.
[0101] In this embodiment, an example has been shown in which processing time is used to grasp the occurrence of a problem, but the present invention is not limited to this, and data flow rate (described later) or the like may also be used.
[0102] In this embodiment, the trace data may have a data type and an event type. Also, the trace data may be recorded for each data type or each event type.
[0103] <Second Example> A control method for a computer system according to a second embodiment of the present invention will be described with reference to Fig. 5. In this embodiment, in the computer system, quality management (state management / health check) of the system is performed using trace units 14_1 to 14_N of operation units 11_1 to 11_N.
[0104] In the computer system 10, for example, the tracing unit 14_1 of the calculation unit 11_1 includes a plurality of event generators 16_1_1 to 16_2_N.
[0105] FIG. 5 is a flowchart illustrating a control method for the computer system 10 according to this embodiment.
[0106] First, for example, the calculation unit 11_1 of the computer system 10 uses a plurality of event generators 16_1_1 to 16_2_N to collect the time when data passes through each of the event generators 16_1_1 to 16_2_N (step S21).
[0107] In detail, the calculation unit 11_1 collects timestamp values (e.g., a first timestamp value and a second timestamp value) based on the detection of a predetermined event in each of the event generators (e.g., the event generator 16_1_1 and the event generator 16_2_1) located at different locations.
[0108] Next, the time required to pass through the specific section in the calculation unit 11_1, that is, the processing time, is calculated by finding the difference between the collected first timestamp value and second timestamp value (step S22).
[0109] Next, the processing time is compared with a predetermined threshold value that has been set in advance (step S23).
[0110] If the comparison result shows that the processing time is greater than the threshold, the detection of an abnormality in the computer system 10 is notified (step S24).
[0111] In this way, according to the computer system management method of this embodiment, it is possible to monitor whether the computer system is operating normally.
[0112] Furthermore, it is not necessary to perform time analysis processing on all data. For example, by observing the time (processing time) required to pass through a specific section in the calculation units 11_1 to 11_N at a predetermined measurement interval and determining whether the processing time falls within a predetermined range or exceeds a predetermined threshold, it is possible to monitor whether the computer system is operating normally.
[0113] Moreover, test data may be inputted and the time required for this data to pass through the inside of the calculation units 11_1 to 11_N may be analyzed.
[0114] In this embodiment, an example is shown in which the processing time is measured directly using the timestamp value acquired by the event generator, but the processing time may also be measured using the timestamp value recorded in the trace buffer after the timestamp value is recorded in the trace buffer.
[0115] In this embodiment, an example has been shown in which the status of a computer system is grasped using processing time, but this is not limiting, and data flow rate or the like may also be used.
[0116] In this embodiment, an example has been shown in which the calculation unit controls the computer system, but the host unit may also control the computer system.
[0117] In this embodiment, the trace data may include a data type along with a timestamp value. It may also include an instance ID, an event type, and arbitrary data. Furthermore, the trace data may be recorded for each data type or event type.
[0118] <Third Example> A control method for a computer system 10 according to a third embodiment of the present invention will be described with reference to Fig. 6. In this embodiment, in the computer system 10, flow management of the computer system 10 is performed using trace units 14_1 to 14_N of operation units 11_1 to 11_N.
[0119] In the computer system 10, the trace units 14_1 to 14_N of the calculation units 11_1 to 11_N record a timestamp value and an instance ID (information indicating a location where an event is detected) as trace data in the trace buffer 18, and the host unit 12 reads out the trace data.
[0120] FIG. 6 is a flowchart illustrating a control method for the computer system 10 according to this embodiment.
[0121] First, the host unit 12 collects, as trace data from the trace buffer 18, for example, timestamp values and instance IDs (information indicating the location where the event is detected) acquired by the event generator 16_1 at any location for different events (step S31).
[0122] In detail, the event generator 16_1 collects the timestamp values of the beginning and end of passing data, which are acquired upon detection of events, for example, the beginning and end of passing data, and stored in the trace buffer 18.
[0123] Next, the difference between the first timestamp value and the last timestamp value is calculated as the data transit time.
[0124] Next, the amount of data per unit time (data flow rate) at a predetermined location is calculated by dividing the amount of input data (or output data) set in advance by the time it takes for the data to pass (step S32).
[0125] Next, the data flow rate is compared with a predetermined threshold value that has been set in advance (step S33).
[0126] If the comparison result shows that the data flow rate is greater than the threshold, a data flow (route) is set to avoid data concentration. For example, when assigning a route, the route is set by avoiding routes that exceed a predetermined threshold identified by an instance ID (information indicating the location where an event is detected) (step S34).
[0127] In this way, according to the computer system management method of this embodiment, it is possible to set a data flow (path) so as to avoid data concentration.
[0128] In this embodiment, an example has been shown in which the processing time is measured using the timestamp value recorded in the trace buffer, but the processing time may also be measured directly using the timestamp value acquired by the event generator.
[0129] Furthermore, although an example has been given in which the host unit controls the computer system, the computing unit may control the computer system.
[0130] Furthermore, if the trace data includes information about the data type, the host unit 12 can grasp the operating status for each data type.
[0131] As a result, in the computer system 10, when data is concentrated in only a specific flow, data can be migrated from that flow to other flows with lower loads, thereby avoiding data concentration and enabling flow management.
[0132] The trace data may also be recorded by instance ID, event type, or arbitrary data. The trace data may also be recorded by data type or event type.
[0133] Furthermore, if a failure that has occurred in a specific flow or at a specific location is detected, the flow can be managed so that a route that bypasses the flow is set when a failure occurs.
[0134] Furthermore, if a malfunction occurs in a calculation unit having multiple flows with different data paths, flows other than the malfunctioning flow can be diverted to other calculation units, and then the calculation unit can be replaced, reset, analyzed, etc.
[0135] <Example of calculation unit measurement> An example of measuring the processing time and data flow rate in the calculation unit in the control method for the computer system 10 according to this embodiment will be described with reference to FIGS. 7A and 7B.
[0136] In this measurement example, for example, input data is observed by an event generator 16_1_1 in the preceding stage of the arithmetic unit 15_1(1), and input data is observed by an event generator 16_2_1 in the subsequent stage.
[0137] When the head of the input data passes through the event generator 16_1_1, an event occurs and the timestamp value of the head of the input data is acquired by the event generator 16_1_1.
[0138] Similarly, when the end of the input data passes through the event generator 16_1_1, an event is triggered, and the timestamp value of the end of the input data is acquired by the event generator 16_1_1.
[0139] On the other hand, when the head of the output data passes through the event generator 16_2_1, an event occurs and the timestamp value of the head of the output data is acquired by the event generator 16_2_1.
[0140] Similarly, when the end of the output data passes through the event generator 16_2_1, an event is triggered, and the timestamp value of the end of the output data is acquired by the event generator 16_2_1.
[0141] An example of trace data is shown in Fig. 7A, which shows a timestamp value (Timestamp:Dec, Timestamp:0x), an instance ID (Ins), an event ID (Evt), a decoded event ID (Dec), a TID, and event data (EventData).
[0142] In the decoded event ID (Dec), H indicates the beginning of the data and L indicates the end of the data.
[0143] The amount of input data and output data is 1 MB. The operating frequency of the arithmetic units 11_1 to 11_N is 250 MHz (4 ns / cycle).
[0144] The timestamp value (Timestamp:Dec) at the beginning of the input data is "406514" and the instance ID (Ins) is "10" (upper row within the dotted box 41). Also, the event ID (Dec) "HR-" indicates the passage of the beginning of the data as an event occurrence.
[0145] Similarly, the timestamp value (Timestamp:Dec) at the beginning of the output data is "401656" and the instance ID (Ins) is "11" (lower row within dotted box 41). Also, the event ID (Dec) "HR-" indicates the passing of the beginning of the data as an event occurrence.
[0146] The timestamp value (Timestamp:Dec) of the end of the input data is "547791" and the instance ID (Ins) is "10" (upper row in dotted box 42). The event ID (Dec) "-LR-" indicates the passing of the end of the data as an event occurrence.
[0147] Similarly, the timestamp value (Timestamp:Dec) at the beginning of the output data is "547794" and the instance ID (Ins) is "11" (lower row within dotted box 42). Also, the event ID (Dec) "-LR-" indicates the passing of the end of the data as an event occurrence.
[0148] 7B shows a schematic diagram of the relationship between input data 43 and output data 44. The time it takes for input data 43 to pass through event generator 16_1_1 is calculated as 146227 cycles = 584.9 μsec from the difference (arrow 45) between the last timestamp value (547791 cycles) and the first timestamp value (401564 cycles). Therefore, the input throughput, i.e., the data flow rate, is calculated as 1MB / 584.9 μsec = approximately 1.8GB / sec.
[0149] Similarly, the input throughput, ie, the data flow rate, can be calculated from the difference (arrow 46) between the timestamp value at the end of the output data 44 and the timestamp value at the beginning.
[0150] In this way, in the calculation unit, the data flow rate is obtained by dividing the amount of data by the difference between the beginning and end of the timestamp value of the data.
[0151] In addition, the processing time (latency) is 92 cycles, calculated from the difference between the timestamp values at the start of output and the start of input, i.e., the difference between the timestamp value at the beginning of output data 44 (401,656 cycles) and the timestamp value at the beginning of input data 43 (401,564 cycles).
[0152] In this way, the processing time of the input data in the calculation unit is obtained from the difference between the timestamp values of the output start and the input start.
[0153] As described above, in the computer system control method according to the present invention, the data processing time and data flow rate can be obtained using the timestamp values of the input data and output data.
[0154] <Effects> According to the computer system and its management method relating to the embodiments and examples of the present invention, when a calculation process in a computer system stops midway or does not complete normally, it is possible to easily detect and identify the calculation unit among multiple calculation units whose processing has stopped.
[0155] Furthermore, since the trace buffer 18 records data, it is possible to resume processing from the middle of the process. As a result, it is not necessary to repeat processing that has already been executed from the beginning. Furthermore, if processing does not complete normally, it is possible to shorten the processing time.
[0156] Furthermore, since the host unit 12 can centrally manage the status of each processing unit, it is possible to set up a data flow that does not pass through a processing unit whose processing has stopped, thereby reducing the number of data whose processing is not completed normally.
[0157] In addition, because events can be detected for each data type (user ID, session ID, stream ID, service ID) and trace data can be accumulated, quality control and defect analysis can be easily performed by focusing on specific data types.
[0158] Furthermore, since the trace units 14_1 to 14_N are provided independently of the calculation units, the abnormal state of the calculation units can be maintained.
[0159] In addition, the flow can be managed at the granularity of each user (each session).
[0160] Furthermore, an event generator that detects an event can be inserted at any point, making it possible to detect defects that occur only in specific flows.
[0161] <Second embodiment> A computer system and a control method thereof according to a second embodiment of the present invention will be described with reference to Fig. 8. A computer system 10 according to this embodiment has a similar configuration to that of the first embodiment.
[0162] In the computer system 10 according to the first embodiment, the trace buffer 18 overflows, making it difficult to record the trace data, so it is necessary to erase the trace data.
[0163] <Control method of computer system> FIG. 8 shows a flowchart of a control method for a computer system according to this embodiment.
[0164] In the computer system according to this embodiment, at least the data type is recorded as trace data together with a timestamp value in the trace buffer 18. Here, the trace data may include an instance ID, an event type, or any other data. Alternatively, the trace data may be recorded for each data type or each event type.
[0165] First, the host unit 12 monitors the trace buffers 18 of the arithmetic units 11_1 to 11_N at a predetermined cycle (step S51).
[0166] Next, it is determined whether or not a plurality of trace data for the same data type are recorded in the trace buffer 18 (step S52).
[0167] If the determination result shows that multiple pieces of trace data for the same data type are recorded in the trace buffer 18, the trace data with the most recent timestamp value among the multiple pieces of trace data is retained (recorded), and previously recorded trace data (trace data other than the most recent trace data) is erased (step S53).In this way, the detection time of the event recorded in the trace buffer is erased based on the detection time of the event.
[0168] According to the computer system and control method of the present embodiment, even if the storage capacity of the trace buffer 18 is limited, buffer overflow can be suppressed and the physical buffer size can be reduced.
[0169] Furthermore, since there is no need to install an excessive amount of buffer, the power consumption of the calculation unit can be reduced, and power efficiency can be improved.
[0170] In addition, by recording trace data for each data type (user ID, session ID, stream ID, service ID), it is possible to provide a highly flexible computer system, such as by retaining trace data for specific data types (such as services with the highest priority) for a relatively long period of time, thereby increasing reliability.
[0171] Naturally, this embodiment also provides the same effects as the first embodiment.
[0172] <Modification 1 of the Second Embodiment> A computer system and a control method thereof according to a first modification of the second embodiment of the present invention will be described below. A computer system 10 according to this modification has the same configuration as the first embodiment.
[0173] In the computer system according to this modification, at least the data type is recorded as trace data together with a timestamp value in the trace buffer 18. Here, the trace data may include an instance ID, an event type, or any other data. Alternatively, the trace data may be recorded for each data type or each event type.
[0174] In this modified example, before the calculation units 11_1 to 11_N of the computer system 10 record (write) the trace data (latest trace data) received from the event generators 16_1_1 to 16_2_N in the trace buffer 18, they determine whether or not trace data of the same data type as the latest trace data is recorded among the trace data already held in the trace buffer 18.
[0175] If the result of the determination is that trace data of the same data type has been recorded, the data already held (data other than the latest data) is erased, and the latest trace data is subsequently recorded, i.e., the latest trace data is overwritten. In this way, the detection time of the event recorded in the trace buffer is overwritten based on the detection time of the event.
[0176] Here, a flag or the like may be added to prevent overwriting, and whether or not overwriting is permitted may be determined based on the presence or absence of the flag.
[0177] This provides the same effect as the present embodiment.
[0178] <Modification 2 of the Second Embodiment> A computer system and a control method thereof according to a second modification of the second embodiment of the present invention will be described below. A computer system 10 according to this modification has the same configuration as that of the first embodiment.
[0179] In the computer system according to this modification, at least a timestamp value is recorded as trace data in the trace buffer 18. Here, the trace data may include a data type, an instance ID, an event type, or any other data. Alternatively, the trace data may be recorded for each data type or event type.
[0180] In this modification, first, the data processed by the arithmetic unit 15_1(1) of the arithmetic unit 11_1 at the front stage is transmitted to the arithmetic unit 15_1(2) of the arithmetic unit 11_2 at the rear stage.
[0181] Next, the arithmetic section 11_2 at the subsequent stage sends a reception completion notification to the arithmetic section 11_1 at the preceding stage.
[0182] Finally, when the front-stage calculation unit 11_1 receives a reception completion notification from the rear-stage calculation unit 11_2, it erases the data in the trace buffer 18 of the front-stage calculation unit 11_1.
[0183] Alternatively, the calculation unit 11_2 at the subsequent stage may send a reception completion notification to the host unit 12, and erase the data in the trace buffer 18 of the calculation unit 11_1 at the preceding stage in response to an instruction from the host unit 12.
[0184] This provides the same effect as the present embodiment.
[0185] Furthermore, in this embodiment, when the host unit 12, which has a larger storage capacity than the trace buffer 18, retrieves and records data from the trace buffer 18, the data from the trace buffer 18 may be erased when the host unit 12 retrieves the data or when the retrieval is completed.
[0186] In this embodiment, the data in the trace buffer 18 may be erased when a preset time has elapsed.
[0187] <Third embodiment> A computer system and a control method thereof according to a third embodiment of the present invention will now be described. A computer system 10 according to this embodiment has the same configuration as the first embodiment.
[0188] In the computer system 10 according to the first embodiment, multiple processing units operate in conjunction with each other, making it difficult to synchronize the times of the timestamps measured by the multiple processing units.
[0189] In particular, when the operating frequencies of the respective arithmetic units are different, the values of the clock counters based on the operating frequencies are different, making it difficult to synchronize the times of the time stamps.
[0190] <Control method of computer system> In the computer system and control method thereof according to this embodiment, the host unit 12 centrally sets the timestamp value to a predetermined value, for example, an initial value. Fig. 9 shows a flowchart of the control method of the computer system according to this embodiment.
[0191] First, the host unit 12 sets a predetermined value as the initial value of the timestamp value (step S61).
[0192] Here, the timestamp value when the calculation units 11_1 to 11_N start measuring the input data may be set as the initial value.
[0193] Alternatively, the host unit 12 may write the start of counting to the timestamp unit 17, and the timestamp value at the time when counting starts may be set as the initial value.
[0194] Next, each of the calculation units 11_1 to 11_N acquires a timestamp value (step S62).
[0195] Next, the host unit 12 compares the acquired timestamp value with a predetermined reference value for the timestamp value (step S63).
[0196] Here, a predetermined reference value for the timestamp value is set in advance, and indicates a predetermined value for the timestamp value, for example, an allowable range of deviation from an initial value.
[0197] Finally, if the timestamp value of each of the calculation units 11_1 to 11_N exceeds the allowable range of a predetermined reference value, the host unit 12 resets the timestamp value of each of the calculation units 11_1 to 11_N to a predetermined value, for example, an initial value (step S64).
[0198] According to the computer system and its control method of this embodiment, when operation units with different operating frequencies coexist (for example, when the operating frequencies differ depending on the operation content of the operation units), the timestamp values can be handled uniformly, making it easy to analyze the trace data after it has been collected.
[0199] Furthermore, since it becomes possible to easily synchronize a plurality of processing units, scalability can be improved.
[0200] Furthermore, since the host unit 12 centrally sets the timestamp value to a predetermined value, for example, an initial value, when the measurement starts, the deviation of the clock counters between the calculation units can be easily corrected.
[0201] Naturally, this embodiment also provides the same effects as the first embodiment.
[0202] <Modification 1 of the third embodiment> A computer system and a control method thereof according to a first modification of the third embodiment of the present invention will be described below. A computer system 10 according to this modification has the same configuration as the first embodiment.
[0203] In this modification, when the frequencies differ among the arithmetic units 11_1 to 11_N, the host unit 12 adjusts the difference in the operating frequencies among the arithmetic units 11_1 to 11_N.
[0204] For example, if the frequency of the calculation unit 11_1 is 100 MHz, one clock cycle is 10 nanoseconds, and the frequency of the calculation unit 11_2 is 200 MHz, one clock cycle is 5 nanoseconds, the host unit 12 doubles the timestamp value of the calculation unit 11_1, thereby synchronizing the timestamp values of the calculation unit 11_1 and the calculation unit 11_2.
[0205] Alternatively, the host unit 12 may halve the timestamp value of the calculation unit 11_2, thereby synchronizing the timestamp values of the calculation unit 11_1 and the calculation unit 11_2.
[0206] Furthermore, after the host unit 12 reads out the trace data, a conversion reference value (for example, 100 MHz) may be set and the counter values of all the calculation units may be converted to match the reference value 100 MHz.
[0207] As described above, in the computer system and control method according to this modification, the host unit multiplies the timestamp value of one of the processing units (e.g., processing unit 11_2) by a coefficient set in the other processing unit (e.g., processing unit 11_2) so that the operating frequency of the other processing unit (e.g., processing unit 11_2) is the same as the operating frequency of the other processing unit (e.g., processing unit 11_1). Here, there may be multiple other processing units. In this way, the difference between the counter values that differ for each processing unit is adjusted.
[0208] As a result, the counter values of the respective calculation units become the same, and the same effect as that of this embodiment is achieved.
[0209] <Modification 2 of the Third Embodiment> A computer system and a control method thereof according to a second modification of the third embodiment of the present invention will be described below. A computer system 10 according to this modification has the same configuration as that of the first embodiment.
[0210] In this modification, when the frequencies differ among the calculation units 11_1 to 11_N, the difference in operating frequency is adjusted among the calculation units 11_1 to 11_N.
[0211] First, a different conversion value (coefficient) is set in advance for each arithmetic unit according to the frequency. For example, when the frequency of the arithmetic unit 11_1 is 100 MHz and the frequency of the arithmetic unit 11_2 is 200 MHz, if the reference value is set to 100 MHz, the conversion value (coefficient) of the arithmetic unit 11_1 will be "1" and the conversion value (coefficient) of the arithmetic unit 11_2 will be "1 / 2".
[0212] Next, before recording the timestamp value in the trace buffer 18, the value obtained by multiplying the timestamp value by a conversion value (coefficient) is recorded.
[0213] Furthermore, after the timestamp value is recorded in the trace buffer 18, the clock counter value may be converted in the same manner when it is read out.
[0214] As described above, in the computer system and control method according to this modification, in the preceding stage of the trace buffer of one of the processing units (e.g., processing unit 11_1), a coefficient set in the other processing unit (e.g., processing unit 11_2) is multiplied by the timestamp value of the other processing unit (e.g., processing unit 11_2) so that the operating frequency of the other processing unit (e.g., processing unit 11_2) is the same as the operating frequency of the other processing unit (e.g., processing unit 11_1). Here, there may be a plurality of other processing units. In this way, the difference between the counter values that differ for each processing unit is adjusted.
[0215] As a result, the counter values of the respective calculation units become the same, and the same effect as that of this embodiment is achieved.
[0216] In an embodiment of the present invention, when the calculation unit performs measurements of processing time, data flow rate, etc., or judgments of defects, etc., the calculation unit may perform these measurements using a calculator, or separate processing functions for measurement, judgment, etc. may be provided within the calculation unit.
[0217] In the embodiment of the present invention, further effects can be achieved by combining the first embodiment, the second embodiment, and the third embodiment.
[0218] In the embodiments of the present invention, examples of the structure, dimensions, materials, etc. of each component in the configuration and management method of the computer system are shown, but the present invention is not limited to these. Anything that can demonstrate the functions and effects of the computer system can be used. [Industrial Applicability]
[0219] The present invention can be applied to computer systems in the field of information processing. [Explanation of symbols]
[0220] 10. Computer Systems 11_1~11_N Arithmetic unit 12 Host Club 13 Internal Communications Department 14_1~14_N Trace section
Claims
1. A plurality of calculation units are provided, Data is transferred between the plurality of calculation units, the calculation unit includes a trace unit that, upon detection of a predetermined event, records trace data including a timestamp value that is a detection time of the event, The trace data includes information specifying a location where the event is detected. A computer system comprising:
2. 1. A computer system for processing input data, comprising: A plurality of calculation units; a host unit connected to the plurality of calculation units and controlling the plurality of calculation units; Equipped with The processed data is transferred between the plurality of calculation units, a trace unit that records trace data when the calculation unit detects a predetermined event from the input data; Equipped with The trace data includes a timestamp value that is the time when the event was detected and information that specifies the location where the event was detected. A computer system comprising:
3. The trace data further includes the type of the input data.
3. The computer system of claim 2.
4. The trace data further includes information for distinguishing the content of the event.
4. The computer system according to claim 2 or 3.
5. The trace data further includes service priority information.
5. A computer system according to claim 2, wherein the computer system comprises:
6. The trace portion is a time stamp unit including at least one clock counter, which counts time based on the operating frequency of the calculation unit; an event generator that is inserted into an arbitrary location of the calculation unit, detects the event, and acquires the timestamp value upon detection of the event; at least one trace buffer that records the trace data in response to detection of the event; 6. A computer system according to claim 2, comprising:
7. the calculation unit includes a calculator, The event generator is newly added to at least one of the upstream and downstream stages of the computing unit, and the computing unit is added.
7. The computer system of claim 6.
8. the calculation unit includes a plurality of calculation units, The event generator in at least one of the preceding stage and the succeeding stage of one of the plurality of computing units is deleted, and the one computing unit is deleted.
7. The computer system of claim 6.
9. one of the event generators that detects one of the events and acquires one of the timestamp values upon detection of the one of the events; another event generator that detects the other event and acquires the other timestamp value in response to the detection of the other event; Either the host unit or the calculation unit measures the difference between the one timestamp value and the other timestamp value.
9. A computer system according to any one of claims 6 to 8.
10. the event generator detects the beginning and end of the input data and obtains the beginning timestamp value and the ending timestamp value; Either the host unit or the calculation unit measures the data flow rate by dividing the amount of input data by the difference between the first timestamp value and the last timestamp value.
10. A computer system according to any one of claims 6 to 9.
11. the trace buffer records at least the timestamp value and information indicating a location where the event is detected in the trace data; When a malfunction occurs in either the host unit or the arithmetic unit, the processing is resumed based on the trace data and using any trace data from a stage before the location where the malfunction occurred in the arithmetic unit. A computer system according to any one of claims 6 to 10.
12. A control method for a computer system comprising a plurality of processing units and a host unit, wherein the plurality of processing units each comprise an event generator, a time stamp unit, and a trace buffer, and data is input to the processing units, the method comprising: a step in which the time stamp unit counts time based on the operating frequency of the calculation unit; the event generator detecting a predetermined event from the data, and acquiring a timestamp value and information specifying a location where the event was detected upon detection of the event; A control method for a computer system comprising:
13. an event generator detecting an event from the data, and acquiring a timestamp value in response to the detection of the event; a step of another event generator detecting another event from the input data and acquiring another timestamp value in response to the detection of the other event; a step in which either the host unit or the calculation unit calculates a processing time from the difference between the one timestamp value and the other timestamp value; The method for controlling a computer system according to claim 12, comprising:
14. the event generator detecting a beginning of the input data, and acquiring a timestamp value of the beginning in response to the detection of the beginning; the event generator detecting the end of the input data, and acquiring a timestamp value of the end upon detecting the end; calculating a difference between the first timestamp value and the last timestamp value as a transit time of the data; a step in which either the host unit or the calculation unit calculates a data flow rate by dividing the amount of input data by the transit time; The method for controlling a computer system according to claim 13, comprising:
15. the trace buffer recording any data; a step in which either the host unit or the calculation unit compares a measured value of either the processing time or the data flow rate with a predetermined threshold; a step in which either the host unit or the calculation unit determines that a malfunction has occurred when the measurement value exceeds the predetermined threshold; determining, by one of the host unit and the calculation unit, a location where the malfunction was detected based on information indicating a location where the event was detected; a step in which either the host unit or the arithmetic unit resumes processing using the arbitrary data recorded in the trace buffer of the arithmetic unit located before the point where the malfunction was detected; 15. The method for controlling a computer system according to claim 14, comprising:
16. a step in which either the host unit or the calculation unit compares a measured value of either the processing time or the data flow rate with a predetermined threshold; a step in which either the host unit or the calculation unit notifies the computer system of an abnormality detection when the measurement value exceeds the predetermined threshold value; 15. The method for controlling a computer system according to claim 14, comprising:
17. recording information indicative of where the event is detected in the trace buffer; a step in which either the host unit or the calculation unit compares the data flow rate with a predetermined threshold; a step in which either the host unit or the calculation unit sets a route through which the data passes so as to avoid concentration of the data when the data flow rate exceeds the threshold; 15. The method for controlling a computer system according to claim 14, comprising:
18. In a computer system, a computing device connected to a host unit and controlled by the host unit, Data is transferred from other computing devices, a trace unit that, upon detection of a predetermined event, records trace data including a timestamp value that is the detection time of the event; The trace data includes information specifying a location where the event is detected. A computing device characterized by:
Citation Information
Patent Citations
Computer system and debugging method
JP1998003403A
Network system
JP2001168898A
Distributed system
JP2001243093A
Virtual computer system, monitoring method of virtual computer system and network system
JP2011258098A
Trace messaging device and methods thereof
US20120331354A1