Exception handling method and exception handling module

Through the exception processing module, the target thread is determined and the abort control signal is sent, which solves the problem of low exception processing efficiency in multi-threaded environments and achieves efficient and accurate exception handling.

CN120086046APending Publication Date: 2025-06-03KUNLUNXIN TECHNOLOGY (BEIJING) CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510114589.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

In multi-threaded parallel processing scenarios, how to efficiently and accurately handle the thread where the processor or core that occurs without affecting normal threads has become a challenge.

Method used

Through the exception processing module, the target thread where the exception processor or core is located is determined, and an abort control signal is sent to the processing components occupied by the thread to ensure that the target thread completes exception processing after exiting.

Benefits of technology

It realizes that only exception threads are processed in a multi-threaded environment, avoid affecting other normal threads, and improves the efficiency and accuracy of exception handling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086046A_ABST
    Figure CN120086046A_ABST
Patent Text Reader

Abstract

The invention provides an exception handling method and an exception handling module, and relates to the technical field of processors, in particular to the technical field of processor exception handling. According to the specific implementation scheme, under the condition that a first processing component in a plurality of processing components is abnormal, a target thread where the first processing component is located is determined, and the target thread is one of a plurality of parallel threads; at least one second processing component occupied by the target thread is determined, and the at least one second processing component comprises the first processing component; and sending a pause control signal to the at least one second processing unit, the pause control signal being used for the at least one second processing unit to pause the operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of processors, and particularly to the technical field of processor exception handling. Background Art

[0002] In the prior art, when an exception occurs in a processor in a processor cluster or in a core in a processor, the driver software needs to collect the exception information of the processor or the core through the interface of the processor, and perform exception handling based on the collected exception information. With the development of technology, more and more multi-threaded parallel methods are used in processor clusters or processors for multi-task processing. However, in such a multi-threaded parallel scenario, how to ensure efficient and accurate exception handling for the thread where the processor or core with an exception occurs without affecting normal threads has become a problem to be solved. Summary of the Invention

[0003] The present disclosure provides an exception handling method, an exception handling module, and a storage medium.

[0004] According to one aspect of the present disclosure, there is provided an exception handling method applied to an exception handling module, including:

[0005] When an exception occurs in a first processing component among a plurality of processing components, determining a target thread where the first processing component is located, where the target thread is one of a plurality of parallel threads;

[0006] Determining at least one second processing component occupied by the target thread, where the at least one second processing component includes the first processing component;

[0007] Sending an abort control signal to the at least one second processing component, where the abort control signal is used for the at least one second processing component to abort the operation.

[0008] According to one aspect of the present disclosure, there is provided an exception handling module, including:

[0009] A thread determination unit, configured to determine a target thread where a first processing component is located when an exception occurs in the first processing component among a plurality of processing components, where the target thread is one of a plurality of parallel threads; and determine at least one second processing component occupied by the target thread, where the at least one second processing component includes the first processing component;

[0010] A communication unit, configured to send an abort control signal to the at least one second processing component, where the abort control signal is used for the at least one second processing component to abort the operation.

[0011] According to another aspect of the present disclosure, there is provided an exception handling module, including:

[0012] at least one processor; and

[0013] a memory communicatively connected to the at least one processor; wherein,

[0014] the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute any of the methods in the embodiments of the present disclosure.

[0015] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute any of the methods in the embodiments of the present disclosure.

[0016] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program which, when executed by a processor, implements any of the methods in the embodiments of the present disclosure.

[0017] By adopting the above implementation manners, when an exception occurs in a first processing component among multiple processing components, the exception handling module can determine the target thread where the first processing component is located, send an abort control signal to at least one second processing component occupied by the target thread, and determine that the exception handling of the target thread is completed when it is determined that the target thread exits. In this way, when an exception occurs in one thread among multiple parallel threads, only the exception thread can be processed, so as to accurately solve the exception problem of the thread where the processing component with the exception occurs without affecting other normal threads; in addition, only through the exception handling module for exception handling, the problem of low efficiency caused by the software accessing the hardware interface multiple times to handle exceptions can be avoided, and the efficiency of exception handling is improved.

[0018] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0020] Figure 1 is a flowchart of an exception handling method according to an embodiment of the present disclosure;

[0021] Figure 2 is a scenario diagram of an exception handling module according to an embodiment of the present disclosure;

[0022] Figure 3 It is a schematic diagram of the scenario of an exception record register according to an embodiment of the present disclosure;

[0023] Figure 4 It is a schematic diagram of the scenario of a thread record register according to an embodiment of the present disclosure;

[0024] Figure 5 It is a schematic flowchart of an exception handling method according to an embodiment of the present disclosure;

[0025] Figure 6 It is another schematic flowchart of an exception handling method according to an embodiment of the present disclosure;

[0026] Figure 7 It is a schematic diagram of the connection relationship between the driver layer and the hardware according to an embodiment of the present disclosure;

[0027] Figure 8 It is a schematic block diagram of an exception handling module according to an embodiment of the present disclosure;

[0028] Figure 9 It is a schematic block diagram of an exception handling module according to another embodiment of the present disclosure;

[0029] Figure 10 It is a block diagram for implementing the exception handling module of the embodiment of the present disclosure. Detailed implementation manners

[0030] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0031] Figure 1 It is a schematic flowchart of an exception handling method proposed by an embodiment of the present disclosure, including:

[0032] S110, in the case where an exception occurs in a first processing component among a plurality of processing components, determine the target thread where the first processing component is located, where the target thread is one of a plurality of parallel threads;

[0033] S120, determine at least one second processing component occupied by the target thread, where the at least one second processing component includes the first processing component;

[0034] S130, send an abort control signal to the at least one second processing component, where the abort control signal is used for the at least one second processing component to abort the operation.

[0035] The exception handling method of the embodiments of the present disclosure can be executed by an exception handling module. The exception handling module at least includes an exception handling state machine, which is used to control the execution of the exception handling method provided by the embodiments of the present application.

[0036] Preferably, the exception handling module may be a chip capable of communicating with a processing component. Here, the specific type or name of the chip is not limited.

[0037] Optionally, the exception handling module may also be other hardware modules with computing capabilities and capable of communicating with a processor or a core in the processor.

[0038] It should be understood that the above is only an exemplary description of the exception handling module, and all possible hardware types of the exception handling module are not limited or exhausted here. As long as the hardware can execute the exception handling method provided in this embodiment, it is within the protection scope of this embodiment.

[0039] By adopting the above implementation manner, when an exception occurs in the first processing component among multiple processing components, the exception handling module can determine the target thread where the first processing component is located, send an abort control signal to at least one second processing component occupied by the target thread, and when it is determined that the target thread exits, determine that the exception handling of the target thread is completed. In this way, when an exception occurs in one thread among multiple parallel threads, only the exception thread is processed, so as to ensure that the exception problem of the thread where the faulty processing component is located is accurately solved without affecting other normal threads; in addition, since the exception handling is only performed by the exception handling module without the participation of the driver layer software, the problem of low efficiency caused by the software accessing the hardware interface multiple times to clear the exception is avoided, and the efficiency of exception handling is improved.

[0040] Each thread among multiple parallel threads occupies at least one processing component among multiple processing components, and different threads occupy different processing components.

[0041] Multiple parallel threads can be used to execute multiple tasks, and different threads are used to execute different tasks. Here, the multiple tasks can be multiple tasks belonging to one total task or multiple tasks that are not related to each other. The present embodiment does not limit the relationship between the multiple tasks. The task executed by any one thread can be any one of a computing task, a processing task, an analysis task, a memory access task, etc. The specific content of the tasks executed by each thread is not limited in this embodiment. For the sake of brevity hereinafter, the task executed by a thread is simply referred to as the task corresponding to the thread, and no repeated explanation will be given hereinafter.

[0042] Furthermore, in at least one processing component occupied by the same thread, each processing component can be used to execute or process the subtask corresponding to the processing component in the task corresponding to the thread, and different processing components execute or process different subtasks. The subtask executed by any one processing component can be at least part of any one of computing tasks, processing tasks, analysis tasks, memory access tasks, etc. For the sake of brevity hereinafter, the subtask corresponding to the processing component in the task corresponding to a thread executed by the processing component is referred to as the subtask corresponding to the processing component.

[0043] The processing component includes one of the following: a core in a target processor; a processor.

[0044] In one scenario, multiple processing components can be multiple processors in a processor cluster. That is to say, multiple processors in the processor cluster correspond to one exception handling module.

[0045] The type of any one of the multiple processors can be one of the following: a central processing unit (CPU), a graphics processing unit (GPU), a tensor processing unit (TPU), etc., which are not listed one by one here.

[0046] In this scenario, each thread in multiple parallel threads occupies at least one processor in the processor cluster, and different threads occupy different processors.

[0047] In another scenario, multiple processing components can be multiple cores in a target processor. Among them, the target processor can be any processor with multiple cores, and the type of the target processor can also be any one of CPU, GPU, and TPU. The target processor and the exception handling module are arranged in the same physical device. For example, the physical device can be a server, a laptop, a desktop computer, etc., which are not elaborated one by one here.

[0048] In this scenario, each thread in multiple parallel threads occupies at least one core in the target processor, and different threads occupy different cores.

[0049] In this way, the exception handling method of the present application can perform exception handling on multiple processors in a processor cluster and can also perform exception handling on multiple cores in a processor. In this way, it can ensure flexible adaptation to multiple scenarios and improve the flexible adaptation degree of multiple scenarios.

[0050] In one embodiment, the method further includes: when an exception record of the first processing component is monitored from the exception record table, determining that the first processing component has an exception.

[0051] As Figure 2 shown, in addition to including the exception handling state machine 2011, the exception handling module 201 may further include an exception record register 2012. The exception record register 2012 may be a device with storage or caching functions provided in the exception handling module. An exception record table may be stored in the exception record register, and the exception record table is used to store or record the exception records of the processing components that have exceptions.

[0052] Taking any processing component with an exception as the first processing component as an example, the manner in which the exception record table records or stores the exception records of the processing components is described as follows: The exception record register receives an exception trigger message sent by the first processing component among multiple processing components; based on the exception trigger message sent by the first processing component, the exception record register records or stores the exception record of the first processing component in the exception record table.

[0053] Combined Figure 3 for exemplary illustration: The multiple processing components include processing component 0 to processing component 3. When processing component 0 has an exception, processing component 0 sends an exception trigger message (i.e., exception reporting) to the exception record register, and the exception record register receives the exception trigger message sent by processing component 0; based on the exception trigger message sent by processing component 0, the exception record register records or stores the exception record of processing component 0 in the exception record table.

[0054] Optionally, the exception trigger message sent by the first processing component to the exception record register may include the identifier of the first processing component.

[0055] Optionally, the exception trigger message sent by the first processing component to the exception record register may include the identifier of the first processing component and information for indicating the exception.

[0056] Among them, the information for indicating the exception may include at least one of the following: an indication value for indicating the exception, a description information for indicating the exception, and information for indicating the exception type corresponding to the exception.

[0057] For example, the indication value for indicating the exception can be configured according to the actual situation, such as 1 or 0.

[0058] For example, the description information for indicating the exception can also be configured according to the actual situation, such as "exception" or "error", etc.

[0059] For example, the information used to indicate the exception type corresponding to an exception can be an exception type code, etc., which is not limited here.

[0060] In a possible case, only the exception records of the processing components where exceptions occur are stored or recorded in the exception record table. In this case, the identifier of the processing component where an exception occurs can be saved in the exception record table as the exception record of this processing component.

[0061] Correspondingly, in the case of detecting the exception record of the first processing component from the exception record table, determining that the first processing component has an exception can include: the exception processing state machine in the exception processing module periodically monitors the exception record table saved in the exception record register in the exception processing module. When the exception processing state machine detects an exception record containing the identifier of the first processing component in the exception record table in any one period, it is determined that the first processing component has an exception. Wherein, the duration of the period can be configured according to the actual situation. For example, it can be 0.1, or 0.2 seconds, or longer or shorter. This application does not limit it.

[0062] In another possible case, the exception record table can include the identifiers of all processing components and the indication information associated with each processing component's identifier for indicating whether an exception occurs.

[0063] For example, the indication information used to indicate whether an exception occurs can be represented in a value-taking manner. For example, when the indication information used to indicate whether an exception occurs is a first value, it is used to indicate that the corresponding or associated processing component has an exception. When the indication information used to indicate whether an exception occurs is a second value, it is used to indicate that the corresponding or associated processing component has no exception. The first value and the second value are different. For example, the first value can be 1 and the second value can be 0.

[0064] For example, the indication information used to indicate whether an exception occurs can be represented by descriptive information. For example, when the indication information used to indicate whether an exception occurs is "exception", it is used to indicate that the corresponding or associated processing component has an exception. When the indication information used to indicate whether an exception occurs is "normal", it is used to indicate that the corresponding or associated processing component has no exception.

[0065] Correspondingly, in the case of detecting the exception record of the first processing component from the exception record table, determining that the first processing component has an exception can include: the exception processing state machine in the exception processing module periodically monitors the exception record table saved in the exception record register in the exception processing module. When the exception processing state machine detects in any one period that the identifier of the first processing component in the exception record table is associated with the indication information for indicating an exception, it is determined that the first processing component has an exception.

[0066] In this way, by monitoring the exception handling component in the exception record table of the exception record register, the efficiency of monitoring the exception handling component can be improved.

[0067] In one implementation, determining the target thread corresponding to the first processing component includes: sending an exception handling request message to the driver layer software; and when receiving the exception handling start message sent by the driver layer software, looking up the target thread where the first processing component is located from the thread record table.

[0068] The driver layer software is a layer of software located between the operating system and the hardware device, and can also be regarded as a special software program. The driver layer software can be abbreviated as the driver layer, or the software side, or the software layer. Alternatively, the driver layer software can also be alternatively referred to as the runtime driver, or the software of the driver layer, etc. It can be understood that the meanings represented by these names are the same, and no repeated explanations will be given below.

[0069] Sending an exception handling request message to the driver layer software can be: the exception handling state machine in the exception handling module sends an exception handling request message to the driver layer software.

[0070] The content carried in the exception handling request message can include at least one of the following: the identifier of the target thread, the identifier of each of at least one second processing component occupied by the target thread, the indication information for indicating the exception, etc., which are not limited or exhausted here.

[0071] Correspondingly, the processing executed on the driver layer software side includes: the driver layer software receives the exception handling request message sent by the exception handling state machine in the exception handling module; and sends an exception handling start message to the exception handling state machine in the exception handling module. Further, the processing after the driver layer software receives the exception handling request message sent by the exception handling state machine in the exception handling module can also include: stopping sending the subtasks corresponding to each second processing component to each second processing component; and / or stopping sending the tasks corresponding to the target thread to the target thread.

[0072] As Figure 2 shown, in addition to the exception handling state machine 2011 and the exception record register 2012, the exception handling module 201 processing may further include a thread record register 2013. The thread record register 2013 can be a device with storage or caching functions provided in the exception handling module. The thread record register is used to store the thread record table, and the thread record table is used to record the identifiers of at least one thread, the identifiers of at least one processing component occupied by each thread in at least one thread, and the states of at least one processing component occupied by each thread. Among them, the state of the processing component can include the working state or the working end state.

[0073] Specifically, finding the target thread where the first processing component is located in the thread record table may include: the exception handling state machine finds the thread where the identifier of the first processing component is located in the thread record table in the thread record register.

[0074] Combined with Figure 4 Explanation of the thread record table stored in the thread record register: The thread record table records the identifiers of thread 0 and thread 1. The processing components occupied by thread 0 include processing component 0 and processing component 1 (i.e., the identifiers of processing component 0 and processing component 1). The processing components occupied by thread 1 include processing component 2 and processing component 3 (i.e., the identifiers of processing component 2 and processing component 3). It can be seen that when the first processing component is processing component 0, the thread where the first processing component is located is thread 0.

[0075] In this way, before starting exception handling, the exception handling module only needs to interact with the driver layer software once to start the exception handling process, and does not need to communicate with the driver layer software during the subsequent exception handling process, thereby improving the efficiency of exception handling. In addition, finding the target thread where the first processing component is located in the thread record table can improve the efficiency of finding the target thread.

[0076] In one implementation, sending the abort control signal to the at least one second processing component may include: the exception handling state machine in the exception handling module sends the abort control signal to the second interface of each second processing component in the at least one second processing component through the first interface.

[0077] Wherein, the abort control signal is used to notify each second processing component in the at least one second processing component to abort the execution of the task corresponding to the target thread. More specifically, the abort control signal is used to notify each second processing component to abort the subtask corresponding to the second processing component under the task corresponding to the target thread.

[0078] Furthermore, each second processing component occupied by the target thread can be used to execute at least part of the tasks corresponding to the target thread. For the sake of brevity below, at least part of the tasks corresponding to the target thread executed by each second processing component is referred to as the subtask corresponding to each second processing component. It should be noted that the subtasks corresponding to different second processing components are at least partially different.

[0079] Taking any one of the at least one second processing component as the k-th second processing component as an example, the processing on the side of the k-th second processing component may include: the k-th second processing component receives an abort control signal sent by the exception handling state machine through the first interface through the second interface, where k may be a positive integer, and the abort control signal is used for the k-th second processing component to abort the execution of its corresponding subtask; the k-th second processing component aborts the execution of its corresponding subtask.

[0080] As Figure 2 shown, the first interface is the register control interface 2014 of the exception handling module 201.

[0081] When the processing component is a processor, any one of the second processing components refers to the processor, and the second interface of the processor may be the register interface of the processor.

[0082] When the processing component is the core of the target processor, the at least one second processing component refers to at least one core of the target processor, and the second interface of any one core may be the register interface of the core.

[0083] In one implementation, it further includes: sending an exception clearing signal to the first processing component, where the exception clearing signal is used for the first processing component to clear the exception content.

[0084] The sending of the exception clearing signal to the first processing component may be that the exception handling state machine in the exception handling module sends the exception clearing signal to the second interface of the first processing component through the first interface.

[0085] Correspondingly, the operations performed on the side of the first processing component include: the first processing component receives the exception clearing signal sent by the exception handling state machine through the first interface through the second interface; the first processor clears the exception content. Wherein, the first interface and the second interface are the same as those in the above implementation, and will not be elaborated here.

[0086] The first processing component clears the exception content, which may be: the first processing component clears the exception content according to the current exception type. The specific clearing method is not limited in this application. For example: when the current exception type is a page fault, it may be deleting the virtual address corresponding to the current exception type. Another example: when the current exception type is a divide-by-zero fault, it may be deleting the calculation content corresponding to the current exception type.

[0087] Furthermore, the timing for the exception handling state machine in the exception handling module to send an exception clearing signal to the first processing component is specifically after the exception handling state machine in the exception handling module sends an abort control signal to the second interface of each of the at least one second processing component through the first interface, and before the exception handling state machine in the exception handling module determines that the target thread exits.

[0088] Correspondingly, the first processing component will first receive the abort control signal and then receive the exception clearing signal. Therefore, the processing on the side of the first processing component includes: aborting the execution of its corresponding subtask based on the received abort control signal; in the case of completing the aborting of its corresponding subtask, clearing the exception content based on the exception clearing signal. It should be noted that the first processing component may not have completely aborted the execution of its corresponding subtask before receiving the exception clearing signal, or may have already completed aborting its corresponding subtask. This embodiment does not limit it. The reason for such an operation is to prevent the first processing component from continuously reporting the same exception trigger message during the process of clearing the exception due to not stopping working.

[0089] In this way, by sending an exception clearing signal to the first processing component, the first processing component can clear the exception content, so that the first processing component can resume normal work, and further ensure that the first processing component can normally process subsequent tasks.

[0090] In one implementation manner, the method further includes: determining that the exception handling of the target thread is completed when it is monitored that the status of each of the at least one second processing component recorded in the thread record table is the end-of-work status.

[0091] The determining that the exception handling of the target thread is completed when it is monitored that the status of each of the at least one second processing component recorded in the target record table is the end-of-work status includes: the exception handling state machine in the exception handling module periodically monitors the status of each of the at least one second processing component occupied by the target thread in the thread record table; when the status of each second processing component is the end-of-work status, determining that the target thread exits and determining that the exception handling of the target thread is completed.

[0092] Among them, the method for determining the status of each second processing component as the end-of-work status includes: when each second processing component aborts the execution of the task corresponding to the target thread, determining the status of each second processing component as the end-of-work status. Taking the k-th processing component among the multiple second processing components as an example, when each second processing component aborts the execution of the task corresponding to the target thread, determining the status of each second processing component as the end-of-work status includes: when the k-th second processing component aborts the execution of the k-th sub-task among the tasks corresponding to the target thread, determining the status of the k-th second processing component as the end-of-work status.

[0093] In this way, by monitoring the status of each second processing component in the at least one second processing component recorded in the thread record table, it is determined whether the target thread exits. In this way, it is possible to quickly determine whether the target thread exits, improving the efficiency of determining whether the target thread exits.

[0094] In one implementation, it further includes one of the following: sending an exception handling end message to the driver layer software; in response to an exception status query request sent by the driver layer software, sending an exception handling end message to the driver layer software.

[0095] Sending an exception handling end message to the driver layer software can include two scenarios:

[0096] Scenario 1: After determining that the exception handling of the target thread is completed, the exception handling state machine in the exception handling module directly sends an exception handling end message to the driver layer software.

[0097] Among them, the exception handling end message is used to notify the driver layer software that the exception handling of the target thread is completed. Specifically, the exception handling end message can carry indication information indicating the end of exception handling, and the exception handling end message can also carry at least one of the following: the identifier of the target thread, the identifier of each second processing component.

[0098] Correspondingly, the operations performed by the driver layer software include: receiving the exception handling end message sent by the exception handling state machine in the exception handling module and determining that the exception handling of the target thread is completed.

[0099] Furthermore, after the driver layer software determines that the exception handling of the target thread is completed, it can also re-issue the new task corresponding to the target thread; or re-issue the new sub-task corresponding to each second processing component under the new task corresponding to the target thread to each second processing component.

[0100] The new task corresponding to the target thread may be exactly the same as, completely different from, or partially the same as and partially different from the task corresponding to the target thread. Further, the new subtask corresponding to each second processing component may be exactly the same as, completely different from, or partially the same as and partially different from the subtask corresponding to the second processing component.

[0101] Scenario 2: After determining that the exception handling of the target thread is completed, the exception handling state machine in the exception handling module responds to the exception status query request sent by the driver layer software, and the exception handling state machine in the exception handling module sends an exception handling end message to the driver layer software.

[0102] Correspondingly, the operations performed on the driver layer software side may be: periodically sending an exception status query request to the exception handling state machine in the exception handling module; and determining that the exception handling of the target thread is completed when receiving the exception handling end message sent by the exception handling state machine.

[0103] Among them, the exception status query request can be used to query the status of the target thread. The exception status query request may carry the identifier of the target thread and the request content for indicating the query status. The related description of the exception handling end message is the same as that in the foregoing embodiments and will not be elaborated here.

[0104] Among them, the timing when the driver layer software starts to periodically send an exception status query request to the exception handling state machine in the exception handling module may be: after receiving the exception handling request message sent by the exception handling module.

[0105] Further, after the driver layer software determines that the exception handling of the target thread is completed, it may also stop sending an exception status query request for querying the status of the target thread to the exception handling module. And, after the driver layer software determines that the exception handling of the target thread is completed, it may also re-issue the new task corresponding to the target thread; or re-issue the new subtask corresponding to each second processing component under the new task corresponding to the target thread to each second processing component.

[0106] In this way, after determining that the exception handling of the target thread is completed, the driver layer software is notified of the completion of the exception handling of the target thread through the exception handling end message, and then the driver layer software can re-issue the task corresponding to the target thread, ensuring that the task corresponding to the target thread can be normally executed.

[0107] In one implementation manner, the processing performed by the exception handling module further includes: when the exception handling state machine monitors the exception record of the first processing component from the exception record table, the exception handling state machine transitions from the idle state to the ready state.

[0108] In addition, the processing performed by the exception handling module further includes: after the exception handling state machine sends an exception handling end message to the driver layer software, the exception handling state machine transitions from the ready state to the idle state.

[0109] Among them, the exception handling state machine being in the idle state may refer to the exception handling state machine being in a state of only receiving messages, and / or the exception handling state machine being in a state where it can only monitor whether there is an exception record of a processing component in the exception record register of the exception handling module.

[0110] The exception handling state machine being in the ready state may refer to the exception handling state machine being in a state of performing exception handling. The exception handling performed by this exception handling state machine may include each processing performed by the exception handling state machine in the above-mentioned exception handling method provided in this embodiment, which will not be elaborated here.

[0111] Combined Figure 5 , an exemplary description of the above exception handling method is as follows:

[0112] S501, the exception handling state machine in the exception handling module is in the idle state. At this time, the exception handling state machine will continuously monitor whether there is an exception record of any processing component in the exception record table of the exception record register.

[0113] Furthermore, when the exception handling state machine detects the exception record of the first processing component in the exception record table of the exception record register, the exception handling state machine transitions from the idle state to the ready state, and then executes S502 to S505.

[0114] S502, the exception handling state machine sends an exception handling request message to the driver layer software. When receiving the exception handling start message sent by the driver layer software, it searches for the target thread where the first processing component is located in the thread record table.

[0115] S503, the exception handling state machine determines at least one second processing component occupied by the target thread in the thread record table of the thread record register.

[0116] S504, the exception handling state machine sends an abort control signal to the at least one second processing component through the register control interface. Specifically, the exception handling state machine sends an abort control signal to each of the at least one second processing component through the register control interface.

[0117] S505, the exception handling state machine sends an exception clearing signal to the first processing component through the register control interface.

[0118] S506. After the exception handling state machine determines that the exception handling of the target thread is completed (or after sending an exception handling end message to the driver layer software), the exception handling state machine transitions from the ready state to the idle state and returns to S501.

[0119] Combined with Figure 6 , another exemplary description of the exception handling method is as follows:

[0120] S601. When an exception occurs in the first processing component among multiple processing components, the first processing component among the multiple processing components sends an exception trigger message to the exception record register in the exception handling module.

[0121] It should be noted that in Figure 6 , for the sake of simplicity, the processing component is used to represent one or more second processing components, and the first processing component is included therein. The following will not be repeatedly explained.

[0122] S602. When the exception handling state machine in the exception handling module detects the exception record of the first processing component from the exception record table in the exception record register, it determines that the first processing component has an exception. Then, the exception handling state machine sends an exception handling request message to the driver layer software.

[0123] The method of recording or storing the exception records of the processing components in the exception record table includes: the exception record register in the exception handling module receives the exception trigger message sent by the first processing component among the multiple processing components; based on the exception trigger message sent by the first processing component, the exception record register records or stores the exception record of the first processing component in the exception record table.

[0124] S603. When the driver layer software receives the exception handling request message sent by the exception handling state machine in the exception handling module, it determines that a processing component exception has been detected.

[0125] S604. The driver layer software sends an exception handling start message to the exception handling state machine in the exception handling module.

[0126] Optionally, after the driver layer software executes S604, it can periodically send an exception status query request to the exception handling state machine in the exception handling module.

[0127] Optionally, after the driver layer software executes S604, the operations performed by the driver layer software further include: stopping sending the corresponding subtasks of each second processing component to each second processing component; and / or stopping sending the task corresponding to the target thread to the target thread.

[0128] S605. When the exception handling state machine in the exception handling module receives the exception handling start message sent by the driver layer software, perform exception handling on the target thread.

[0129] Specifically, the exception handling state machine performing exception handling on the target thread includes: determining the target thread where the first processing component is located; determining at least one second processing component occupied by the target thread; sending an abort control signal to the at least one second processing component, and sending an exception clearing signal to the first processing component.

[0130] When the exception handling state machine in the exception handling module determines that the state of each second processing component recorded in the thread record register in the thread table is the end-of-work state, it determines that the exception handling of the target thread is completed; then perform one of the following: send an exception handling end message to the driver layer software; in response to the exception status query request sent by the driver layer software, send an exception handling end message to the driver layer software.

[0131] S606. When the driver layer software receives the exception handling end message sent by the exception handling state machine, determine that the monitoring processing result is the end of the exception handling of the target thread.

[0132] S607. The driver layer software re-issues the new task corresponding to the target thread; or, re-issues the new sub-task corresponding to each second processing component under the new task corresponding to the target thread to each second processing component.

[0133] S608. Each second processing component receives the new task and executes the new task.

[0134] Combined with Figure 7 An exemplary description of the connection relationship between the driver layer and the hardware is as follows Figure 7As shown in the figure, on the hardware side, it includes: processor component 701, register interface 704 of the processor component, and exception handling module 602. On the driver layer side, it includes driver layer software. Among them, the connection between processor component 701 and exception handling module 702 is used for processor component 701 to send an exception trigger message to the exception record register in the exception handling module; the connection between exception handling module 702 and the register interface 704 of the processor component is used for the exception handling state machine in the exception handling module to send an abort control signal and an exception clearing signal to the register interface 704 of the processor component through the register control interface; the connection between the register interface 704 of the processor component and processor component 701 is used for processor component 701 to receive the abort control signal and the exception clearing signal through the register interface 704 of the processor component; the connection between exception handling module 702 and driver layer software 703 is used for at least one of the following: the exception handling state machine in exception handling module 702 sends an exception handling request message to driver layer software 703, the exception handling state machine in exception handling module 702 receives the exception handling start message sent by the driver layer software, the exception handling state machine in exception handling module 702 receives the exception status query request sent by driver layer software 703, and the exception handling state machine in exception handling module 702 sends an exception handling end message to driver layer software 703.

[0135] Figure 8 The figure shows a schematic block diagram of an exception handling module provided by an embodiment of the present disclosure. As Figure 8 shown, it includes:

[0136] A thread determination unit 801, configured to determine a target thread where the first processing component is located when an exception occurs in the first processing component among multiple processing components, where the target thread is one of multiple parallel threads; determine at least one second processing component occupied by the target thread, where the at least one second processing component includes the first processing component;

[0137] A communication unit 802, configured to send an abort control signal to the at least one second processing component, where the abort control signal is used for the at least one second processing component to abort operations.

[0138] The communication unit is configured to send an exception clearing signal to the first processing component, where the exception clearing signal is used for the first processing component to clear exception content.

[0139] As Figure 9 shown, the exception handling module further includes: a component determination unit 901, configured to determine that the first processing component has an exception when an exception record of the first processing component is monitored from the exception record table.

[0140] The communication unit is configured to send an exception handling request message to the driver layer software; the thread determination unit is configured to, when receiving an exception handling start message sent by the driver layer software, look up a target thread where the first processing component is located from a thread record table.

[0141] The thread determination unit is configured to determine that the exception handling of the target thread is completed when it monitors that the status of each of the at least one second processing component recorded in the thread record table is an end-of-work status.

[0142] The communication unit is configured to perform one of the following: send an exception handling end message to the driver layer software; in response to an exception status query request sent by the driver layer software, send an exception handling end message to the driver layer software.

[0143] The multiple processing components include one of the following: multiple cores in a target processor; multiple processors.

[0144] For the specific functions and examples of each module and sub-module of the device according to the embodiments of the present disclosure, reference may be made to the relevant descriptions of the corresponding steps in the above method embodiments, which will not be elaborated herein.

[0145] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0146] According to an embodiment of the present disclosure, the present disclosure also provides an exception handling module, a readable storage medium, and a computer program product.

[0147] As Figure 10 shown, the exception handling module 1000 includes a calculation unit 1001, which can execute various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a random access memory (RAM) 1003. In the RAM 1003, various programs and data required for the operation of the exception handling module 1000 can also be stored. The calculation unit 1001, the ROM 1002, and the RAM 1003 are connected to each other through a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.

[0148] Multiple components in the exception handling module 1000 are connected to the I / O interface 1005, including: an input unit 1006; an output unit 1007; a storage unit 1008; and a communication unit 1009. The communication unit 1009 allows the exception handling module 1000 to exchange information / data with other devices.

[0149] The computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 executes the various methods and processes described above. For example, in some embodiments, the above method can be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed onto the exception handling module 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the computing unit 1001, at least one step of the method described above can be executed. Alternatively, in other embodiments, the computing unit 1001 can be configured to execute the above method in any other suitable manner (e.g., by means of firmware).

[0150] The various embodiments of the systems and technologies described above in this article can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs, the one or more computer programs being executable and / or interpretable on a programmable system including at least one programmable processor, the programmable processor being a special or general programmable processor, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0151] The program code for implementing the methods of the present disclosure can be written in any combination of at least one programming language. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.

[0152] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on at least one wire, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0153] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0154] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0155] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, or a server of a distributed system, or a server incorporating a blockchain.

[0156] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.

[0157] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. An exception handling method, applied to an exception handling module, comprising: When an exception occurs in a first processing component among the multiple processing components, determining a target thread where the first processing component is located, wherein the target thread is one of the multiple parallel threads; determining at least one second processing element occupied by the target thread, wherein the at least one second processing element includes the first processing element; A suspend control signal is sent to the at least one second processing component, wherein the suspend control signal is used for the at least one second processing component to suspend operation.

2. The method according to claim 1, further comprising: An abnormality clearing signal is sent to the first processing component, wherein the abnormality clearing signal is used by the first processing component to clear abnormal content.

3. The method according to claim 1 or 2, further comprising: When an abnormal record of the first processing component is monitored from the abnormal record table, it is determined that an abnormality occurs in the first processing component.

4. The method according to claim 1 or 2, wherein: The determining the target thread where the first processing component is located includes: Send an exception handling request message to the driver layer software; When the exception handling start message sent by the driver layer software is received, the target thread where the first processing component is located is searched from the thread record table.

5. The method according to claim 4, further comprising: When it is monitored that the state of each second processing component in the at least one second processing component recorded in the thread record table is a work completion state, it is determined that the exception processing of the target thread is completed.

6. The method according to claim 5, further comprising one of the following: Sending an exception handling end message to the driver layer software; In response to the abnormal status query request sent by the driver layer software, an abnormal processing end message is sent to the driver layer software.

7. The method according to claim 1, wherein: The processing component includes one of the following: a core in a target processor; a processor.

8. An exception handling module, comprising: a thread determination unit, configured to, when an exception occurs in a first processing unit among the plurality of processing units, determine a target thread where the first processing unit is located, wherein the target thread is one of the plurality of parallel threads; and determine at least one second processing unit occupied by the target thread, wherein the at least one second processing unit includes the first processing unit; The communication unit is used to send a stop control signal to the at least one second processing component, wherein the stop control signal is used for the at least one second processing component to stop operation.

9. The abnormality handling module according to claim 8, wherein the communication unit is used to send an abnormality clearing signal to the first processing component, wherein: The abnormality clearing signal is used by the first processing component to clear abnormal content.

10. The exception handling module according to claim 8 or 9, further comprising: The component determination unit is used to determine that an abnormality occurs in the first processing component when an abnormality record of the first processing component is monitored from the abnormality record table.

11. According to the exception handling module according to claim 8 or 9, the communication unit is used to send an exception handling request message to the driver layer software; the thread determination unit is used to search the target thread where the first processing component is located from the thread record table when receiving the exception handling start message sent by the driver layer software.

12. The exception handling module according to claim 11, wherein: The thread determination unit is used to determine that the exception handling of the target thread is completed when monitoring that the state of each second processing component in the at least one second processing component recorded in the thread record table is a work completion state.

13. The exception handling module according to claim 12, wherein: The communication unit is configured to perform one of the following: Sending an exception handling end message to the driver layer software; In response to the abnormal status query request sent by the driver layer software, an abnormal processing end message is sent to the driver layer software.

14. The exception handling module according to claim 8, wherein: The processing component includes one of the following: a core in a target processor; a processor.

15. An exception handling module, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.

16. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-7.

17. A computer program product, comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Multi-process nested concurrence industrial control method based on finite-state machine

    CN122151795A

  • A multi-process nested concurrent industrial control method based on finite state machines

    CN122151795B