A shared DPU-based failure degradation switchover method, device and medium

CN122802342APending Publication Date: 2026-09-22启朔(深圳)科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611094770.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-22
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0005]因此,本发明提供了一种基于共享DPU的故障降级切换方法解决共享DPU共因故障难以准确定位和受影响服务与节点难以准确确定的问题

Benefits of technology

[0036]The beneficial effects of this invention are as follows: by constructing a dependency graph based on shared DPU runtime association data, differential probes are sent to different computing nodes and DPU services, and multiple dependency paths corresponding to abnormal responses are traced back step by step to the common component where the first independent deviation occurred. This enables the association and location of faulty components, affected DPU services, and nodes to be degraded, providing a unified basis for subsequent resource isolation, node degradation, and service recovery. It also avoids including non-faulty services or irrelevant computing nodes in the handling scope, thereby improving the accuracy and continuity of shared DPU fault handling in industrial cloud computing environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122802342A_ABST
    Figure CN122802342A_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on shared DPU's fault degradation switching method, equipment and medium, it is related to industrial cloud computing control field, including, collection shared DPU operating state, node access state and out-of-band management state, and record the DPU service interaction data during normal operation, form shared DPU operation associated data;According to shared DPU operation associated data, construct shared DPU dependency graph, send differential probe to different computing node and different DPU service, compare probe response with normal response, determine the fault component corresponding to multiple abnormal paths along shared DPU dependency graph, and determine affected DPU service and node to be degraded;The application realizes the accuracy and continuity of improving shared DPU fault disposal under industrial cloud computing environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of industrial cloud computing control, and in particular to a fault degradation switching method, device and medium based on a shared DPU. Background Technology

[0002] As industrial cloud computing evolves towards resource pooling, heterogeneous acceleration, and high-density node collaboration, the Data Processing Unit (DPU) is gradually taking on basic functions such as network forwarding, storage virtualization, security offloading, and service orchestration. To reduce hardware redundancy and improve resource utilization, the deployment method of multiple computing nodes sharing DPU services has begun to be applied to cloud data centers. Existing fault handling typically involves collecting DPU operating metrics, node access status, and out-of-band management information, and triggering service switching, node isolation, or backup resource takeover based on heartbeat timeouts, interface anomalies, or single performance thresholds to maintain the basic continuity of computing tasks.

[0003] In a shared DPU architecture, the same faulty component may be located on multiple compute nodes and the dependency paths of multiple DPU services. Existing technologies mostly rely on single-node alarms or single-service response anomalies as the basis for judgment, lacking a correlation analysis mechanism that combines service call relationships, resource bearing relationships, and common characteristics of multiple abnormal paths. When a common cause of failure occurs in the switching channel, virtual function, driver instance, or management link, multiple nodes may exhibit similar but not independent access anomalies. Single-point judgment makes it difficult to distinguish the source of the fault from the affected object, which can easily lead to a lack of consistent basis for determining the scope of degradation, isolated objects, and subsequent recovery objects, thereby affecting the accuracy of shared DPU fault handling in industrial cloud computing environments. Summary of the Invention

[0004] In view of the aforementioned existing problems, the present invention is proposed.

[0005] Therefore, this invention provides a fault degradation and switching method based on shared DPU to solve the problems of difficulty in accurately locating common-cause faults in shared DPU and difficulty in accurately determining affected services and nodes.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0007] In a first aspect, the present invention provides a fault degradation switching method based on a shared DPU, comprising: collecting the shared DPU's operating status, node access status, and out-of-band management status, and recording DPU service interaction data during normal operation to form shared DPU operating association data; constructing a shared DPU dependency graph based on the shared DPU operating association data, sending differential probes to different computing nodes and different DPU services, comparing the probe responses with normal responses, identifying faulty components corresponding to multiple abnormal paths along the shared DPU dependency graph, and identifying affected DPU services and nodes to be degraded; determining the degradation level of each node to be degraded based on the faulty components, and isolating affected nodes. The system impacts the service-bearing resources corresponding to the DPU service and controls the nodes to be degraded to load the rescue operation environment through an independent escape channel. After the nodes to be degraded enter the rescue operation state, the faulty components are repaired, and the restored DPU service is loaded into a candidate recovery environment isolated from production business, thus obtaining a shared DPU candidate recovery state. Based on the shared DPU candidate recovery state, when the replay response of the DPU service meets the normal response consistency condition and the completion status of the boundary request is determined, a valid service recovery token bound to the DPU service is generated, and the compute nodes are controlled to restore the shared DPU service path according to the service recovery token, thus obtaining shared DPU fault degradation switching information.

[0008] As a preferred embodiment of the fault degradation and switching method based on shared DPU described in this invention, the specific steps for collecting the shared DPU operating status, node access status, and out-of-band management status, and recording DPU service interaction data during normal operation to form shared DPU operating association data are as follows:

[0009] According to the collection cycle, acquire the shared DPU running status, node access status and out-of-band management status, and write the associated information for each status data;

[0010] Retrieve DPU service interaction records during normal operation, and establish associations between service interaction records and corresponding status data based on the association information to form shared DPU operation association data.

[0011] As a preferred embodiment of the fault degradation and switching method based on shared DPU described in this invention, the specific steps of constructing a shared DPU dependency graph based on shared DPU runtime association data and sending differential probes to different computing nodes and different DPU services are as follows:

[0012] Extract the calling and carrying relationships between objects from the shared DPU runtime association data, and construct a shared DPU dependency graph with objects as graph nodes and object relationships as graph edges.

[0013] Select compute nodes and DPU services as probe targets along the shared DPU dependency graph, send differential probe requests according to different probe paths, and record the corresponding probe responses.

[0014] As a preferred embodiment of the fault degradation and switching method based on shared DPU described in this invention, the steps of comparing probe responses with normal responses, determining the faulty components corresponding to multiple abnormal paths along the shared DPU dependency graph, and determining the affected DPU services and nodes to be degraded are as follows:

[0015] Align each probe response with its corresponding normal response, compare response characteristics, retain probe responses that deviate from the normal response range, and mark the dependency links between the probe sending position and the abnormal response position along the shared DPU dependency graph to form an abnormal path;

[0016] The system aggregates abnormal paths corresponding to different compute nodes and DPU services, traces back level by level along the dependency direction, and filters out common components that appear repeatedly in multiple abnormal paths and are located upstream of abnormal forks. It compares the response deviations of common components with those of upstream components. When the upstream component responds normally and the common component is the first to show a response deviation, the common component is identified as the faulty component. When the upstream component also shows a response deviation, the system continues to trace back upstream of the dependency until the faulty component that first showed an independent response deviation is identified.

[0017] Starting from the faulty component, expand along the service dependency direction, extract the service path containing the faulty component, identify the DPU services with abnormal responses in the service path as the affected DPU services, and identify the compute nodes that depend on the affected DPU services and cannot switch to the normal service path as nodes to be degraded.

[0018] As a preferred embodiment of the fault degradation and switching method based on shared DPU described in this invention, the steps of determining the degradation level of each node to be degraded based on the faulty component, isolating the service-bearing resources corresponding to the affected DPU service, and controlling the node to be degraded to load the rescue operation environment through an independent escape channel are as follows:

[0019] Based on the abnormal path corresponding to the faulty component, determine the fault impact range and service replacement capability of each node to be downgraded, and match the corresponding downgrade level according to the downgrade rules;

[0020] Disconnect the access path between the affected DPU service and the production business according to the downgrade level, lock the corresponding service carrying resources, and retain the running path of the unaffected service;

[0021] A rescue activation command is sent to the node to be downgraded via an independent escape route, controlling the node to load the pre-set rescue operation environment and switch to the rescue operation state corresponding to the downgrade level.

[0022] As a preferred embodiment of the fault degradation and switching method based on shared DPU described in this invention, the steps are as follows: After the node to be degraded enters the rescue operation state, the faulty components are repaired, and the recovered DPU service is loaded into a candidate recovery environment isolated from production business to obtain a shared DPU candidate recovery state.

[0023] After the node to be downgraded enters the rescue operation state, the repair strategy corresponding to the faulty component is adopted to stop the faulty component from continuing to carry production requests, clear the abnormal operation state, and perform component recovery operation to restore the faulty component.

[0024] Send a management probe request to the recovered faulty component, verify the operating conditions of the faulty component based on the probe response, and mark the faulty component as recoverable when the faulty component meets the DPU service loading requirements;

[0025] After a faulty component is marked as recoverable, a candidate recovery environment is created in a resource area isolated from production operations. The recovered DPU service and its running configuration are loaded into the candidate recovery environment, the dependent connections corresponding to the DPU service are restored, and the production access entry is closed. When the DPU service is started and can receive verification requests, a shared DPU candidate recovery state is formed.

[0026] As a preferred embodiment of the fault degradation switching method based on shared DPU described in this invention, the step of generating a valid service recovery token bound to the DPU service based on the shared DPU candidate recovery status, when the replay response of the DPU service meets the normal response consistency condition and the completion status of the boundary request is determined, is as follows:

[0027] Retrieve loaded DPU services from the shared DPU candidate recovery status, select the service requests corresponding to the normal operation period, and replay the service requests to the candidate recovery environment according to the original call relationship; align the replay response with the corresponding normal response, compare the response characteristics, and write a replay consistency flag for the corresponding DPU service when the response difference is within the preset consistency range.

[0028] Filter out boundary requests that did not return explicit completion information from the DPU service interaction records before and after the failover, and track the corresponding execution records and status changes based on the request identifier; when the execution record and status change correspond to each other, mark the boundary request as completed; when no execution trace is found and resubmission is allowed, mark the boundary request as incomplete; when there are still conflicts in the completion status, retain the uncertain mark.

[0029] When the DPU service has a replay consistency flag and all corresponding boundary requests have obtained completed or incomplete flags, the recovery verification information corresponding to the DPU service is collected, the recovery verification information is bound to the DPU service, and a valid service recovery token is generated. When the replay response does not meet the normal response consistency conditions or there are boundary requests with uncertain flags, the DPU service is kept in the candidate recovery environment and no valid service recovery token is generated.

[0030] As a preferred embodiment of the fault degradation and switching method based on shared DPU described in this invention, the specific steps for controlling the computing node to restore the shared DPU service path based on the service recovery token to obtain shared DPU fault degradation and switching information are as follows:

[0031] Identify the corresponding DPU service from the service recovery token, determine the compute node that depends on the corresponding DPU service based on the shared DPU dependency graph, and verify whether the service recovery tokens that the compute node depends on are all in a valid state.

[0032] When the recovery conditions are met, the locking of the corresponding service-bearing resources and the isolation of the service path are released, and the access entry of the computing node is switched to the recovered shared DPU service path;

[0033] When service access is normal, the control computing node exits the rescue operation state. When service access is abnormal, the corresponding service recovery token is revoked, and the computing node is switched back to the rescue operation environment. The fault location, node degradation and service recovery information are collected to obtain the shared DPU fault degradation and switching information.

[0034] In a second aspect, the present invention provides a computer device, including a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, it implements any step of the fault degradation switching method based on a shared DPU as described in the first aspect of the present invention.

[0035] Thirdly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the fault degradation switching method based on a shared DPU as described in the first aspect of the present invention.

[0036] The beneficial effects of this invention are as follows: by constructing a dependency graph based on shared DPU runtime association data, differential probes are sent to different computing nodes and DPU services, and multiple dependency paths corresponding to abnormal responses are traced back step by step to the common component where the first independent deviation occurred. This enables the association and location of faulty components, affected DPU services, and nodes to be degraded, providing a unified basis for subsequent resource isolation, node degradation, and service recovery. It also avoids including non-faulty services or irrelevant computing nodes in the handling scope, thereby improving the accuracy and continuity of shared DPU fault handling in industrial cloud computing environments. Attached Figure Description

[0037] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0038] Figure 1 This is a flowchart of a fault degradation switching method based on a shared DPU.

[0039] Figure 2 A flowchart for identifying affected DPU services and nodes to be degraded.

[0040] Figure 3 A flowchart for generating a valid service recovery token.

[0041] Figure 4 A flowchart for generating shared DPU fault degradation switching information. Detailed Implementation

[0042] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0043] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0044] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0045] Reference Figures 1-4As one embodiment of the present invention, this embodiment provides a fault degradation switching method based on a shared DPU, including the following steps:

[0046] S1. Collect the shared DPU running status, node access status, and out-of-band management status, and record the DPU service interaction data during normal operation to form shared DPU running association data.

[0047] According to the collection cycle, the shared DPU running status, node access status, and out-of-band management status are obtained, and the associated information is written for each status data.

[0048] Furthermore, at the arrival of each acquisition cycle, a query request is sent to the shared DPU status interface to receive the running status returned by the shared DPU; access requests are initiated to each node according to the node identifier, and the request return status is recorded; the management status of the corresponding shared DPU is read from the out-of-band management interface; based on the shared DPU, node, and DPU service corresponding to each status data, the shared DPU identifier, node identifier, DPU service identifier, acquisition time, and data source are written to each status data, and the status data within the same acquisition cycle are grouped into the same status record.

[0049] Retrieve DPU service interaction records during normal operation, and establish associations between service interaction records and corresponding status data based on the association information to form shared DPU operation association data.

[0050] Furthermore, request and response records for each DPU service are retrieved from the normal operation logs, and the shared DPU identifier, node identifier, DPU service identifier, and interaction time corresponding to the service interaction records are extracted. The identifier information in the service interaction records is compared item by item with the associated information in the status data, and the status data within the same collection period is matched according to the interaction time. The matched service interaction records are written into the corresponding status data to form shared DPU operation association data.

[0051] It should be noted that the service interaction record includes the requests sent by the compute nodes to the DPU service and the responses returned by the DPU service.

[0052] S2. Construct a shared DPU dependency graph based on the shared DPU runtime association data, send differential probes to different computing nodes and different DPU services, compare the probe responses with normal responses, identify the faulty components corresponding to multiple abnormal paths along the shared DPU dependency graph, and identify the affected DPU services and nodes to be degraded.

[0053] Extract the calling and carrying relationships between objects from the shared DPU runtime association data, and construct a shared DPU dependency graph with objects as graph nodes and object relationships as graph edges.

[0054] Furthermore, node access records, DPU service interaction records, and service bearing records are read from the shared DPU runtime association data, and the corresponding records are merged according to the shared DPU identifier, node identifier, and DPU service identifier; the calling direction between the request sender and receiver is determined based on the corresponding information of DPU service and service bearing resources; the shared DPU, nodes, DPU service, and service bearing resources are written into graph nodes respectively, and graph edges with relationship type and direction are written between graph nodes with calling relationship or bearing relationship to form a shared DPU dependency graph.

[0055] It should be noted that service bearer resources refer to the underlying resources used to load and run DPU services and provide request processing, data storage and communication connections for DPU services, including processing resources, storage resources and communication resources; service bearer records record the binding relationship between DPU services and corresponding service bearer resources through resource identifiers.

[0056] Select compute nodes and DPU services as probe targets along the shared DPU dependency graph, send differential probe requests according to different probe paths, and record the corresponding probe responses.

[0057] Furthermore, along the call direction in the shared DPU dependency graph, different compute nodes and different DPU services are selected as probe objects, and the direct access path and the access path forwarded by associated components are determined from the call edges from the compute node to the corresponding DPU service; a path identifier, a starting node identifier, and a target service identifier are written for each probe path; probe requests with the same probe content and request parameters but different path identifiers are sent to each probe object, the probe responses returned by each probe path are received, and the response content, response status, return time, and graph nodes traversed are recorded according to the probe object identifier and path identifier.

[0058] Align each probe response with its corresponding normal response, compare response characteristics, retain probe responses that deviate from the normal response range, and mark the dependency links between the probe sending location and the abnormal response location along the shared DPU dependency graph to form an abnormal path.

[0059] Furthermore, based on the probe object identifier, path identifier, and request parameters, the corresponding normal response is retrieved from the service interaction records during normal operation, and the probe response is aligned with the normal response item by item according to the response fields; the response status, response content, response duration, and request completion status are compared in turn, and probe responses with inconsistent response status, mismatched key response fields, response duration exceeding the normal range, or incomplete requests are identified as abnormal responses.

[0060] Based on the path identifier corresponding to the abnormal response, locate the probe sending position and the abnormal response position, and trace from the probe sending position to the abnormal response position level by level along the shared DPU dependency graph according to the order in which the probe request passes through the graph nodes; write the traced graph nodes and graph edges into the same abnormal path, and record the probe object, abnormal response characteristics and path identifier to form an abnormal path.

[0061] It should be noted that the normal response range is determined based on historical response samples with identical probe objects, path identifiers, and request parameters during normal operation of the shared DPU. For response status and request completion status, the corresponding normal status in historical normal responses is used as the judgment benchmark. For character-type key response fields, consistency of field content is used as the matching condition. For numerical-type key response fields, the median of historical response samples is used as the benchmark, the absolute deviation of each sample relative to the median is calculated, and the 95th percentile of the absolute deviation is used as the allowable fluctuation range. For response duration, historical response duration samples are arranged in ascending order, and the 95th percentile is used as the upper limit of the response duration. A probe response exceeding the corresponding judgment boundary is determined to be an abnormal response.

[0062] The system aggregates abnormal paths corresponding to different compute nodes and different DPU services, traces back level by level along the dependency direction, and filters out common components that appear repeatedly in multiple abnormal paths and are located upstream of abnormal forks. It compares the response deviations of common components with those of upstream components. When the upstream component responds normally and the common component is the first to show a response deviation, the common component is identified as the faulty component. When the upstream component also shows a response deviation, the system continues to trace back upstream of the dependency until the faulty component that first showed an independent response deviation is identified.

[0063] Furthermore, the abnormal paths are aggregated according to the compute node identifier and DPU service identifier. The graph nodes and edges contained in different abnormal paths are compared item by item, and the number of times each component appears in the abnormal path is counted. The abnormal response location is traced back level by level in the direction of the call source, and the components that appear in multiple abnormal paths at the same time and are located upstream of the fork position of each abnormal path are retained to obtain the common components. The probe responses corresponding to the common components and adjacent upstream components are read according to the dependency order, and the response status, key response fields, response duration and request completion status are compared with the normal response range.

[0064] The response deviation is calculated according to the following formula:

[0065] ;

[0066] in, Representation Component The normalized response deviation is such that the larger the value, the greater the degree to which the component response deviates from the normal state. This indicates the common components currently involved in fault location; The index representing the response characteristic; This represents the total number of numerical response features involved in the deviation calculation; Indicates to component After sending the differential probe, the first [probe name] is extracted from the returned probe response. Numerical response characteristic values; Representation Component During normal operation, the corresponding number The baseline value for each response characteristic is determined by the average or median of similar normal responses; Representation Component Corresponding to the The allowable fluctuation range of each response feature is set based on the fluctuation range during normal operation, the field's allowable error, or the detection accuracy.

[0067] When a common component is the root node of a shared DPU dependency graph and has no adjacent upstream components, the probe response of the common component is compared with the normal response range. If the probe response of the common component exceeds the normal response range, the common component is directly identified as a faulty component. When a common component has adjacent upstream components, if the probe response of the adjacent upstream components is within the normal response range and the probe response of the common component exceeds the normal response range, the common component is identified as a faulty component. When both the common component and the adjacent upstream components have response deviations, the adjacent upstream components are used as new comparison objects, and the process continues to backtrack upwards along the dependency direction until the upstream component responds normally and the current comparison object shows a response deviation. The current comparison object is then identified as the faulty component that first shows an independent response deviation.

[0068] It should be noted that the selection criteria for common components are set based on the number of abnormal paths. The component must appear in at least two abnormal paths with different origins at the same time, and be located before the point where the abnormal path forks.

[0069] Starting from the faulty component, expand along the service dependency direction, extract the service path containing the faulty component, identify the DPU services with abnormal responses in the service path as the affected DPU services, and identify the compute nodes that depend on the affected DPU services and cannot switch to the normal service path as nodes to be degraded.

[0070] Furthermore, starting from the graph node corresponding to the faulty component, the process expands step by step along the dependency edges pointing from the faulty component to the DPU service and compute node, recording the DPU services, service-bearing resources, and compute nodes traversed during the expansion process. The complete dependency links that pass through the faulty component and connect to the compute node are extracted as service paths and saved according to the path identifiers. The probe responses corresponding to the DPU services in each service path are retrieved, and the response status, key response fields, response duration, and request completion status are compared with the normal response range item by item. DPU services that exceed the normal response range are marked as affected DPU services, and the compute nodes that call the affected DPU services are found along the dependency edges.

[0071] For each compute node, search along the shared DPU dependency graph for other service paths that do not pass through the faulty component, and verify the DPU service response and service bearer resource status in the other service paths. If there are other service paths that are complete and respond normally, the compute node is kept in normal operation. If there are no other service paths, or if other service paths contain DPU services with abnormal responses, the compute node is identified as a node to be degraded.

[0072] It should be noted that a normal service path refers to a service path that does not pass through faulty components, has a complete path connection, and where the responses of each DPU service within the path are within the normal response range.

[0073] S3. Determine the degradation level of each node to be degraded based on the faulty component, isolate the service carrying resources corresponding to the affected DPU service, and control the node to be degraded to load the rescue operation environment through an independent escape channel; after the node to be degraded enters the rescue operation state, repair the faulty component, and load the restored DPU service into a candidate recovery environment isolated from the production business to obtain a shared DPU candidate recovery state.

[0074] It should be noted that an independent escape route refers to a management communication link that is independent of the shared DPU service path and can still maintain connectivity when the shared DPU service is abnormal. It is used to transmit rescue initiation commands and return command confirmation information, and specifically consists of at least one of an out-of-band management link and an independent management network.

[0075] Based on the abnormal path corresponding to the faulty component, determine the scope of the fault impact and service replacement capability of each node to be downgraded, and match the corresponding downgrade level according to the downgrade rules.

[0076] Furthermore, based on the abnormal path where the faulty component is located, find the affected DPU services that each node to be degraded depends on, and read the core service identifier from the business deployment configuration to count the number of affected core services; query the backup DPU services, backup service paths and local alternative functions. When at least one alternative method responds normally and the available resources meet the running requirements of the corresponding core service, mark the corresponding core service as replaceable and count the number of replaceable core services.

[0077] When no core services are involved, a mild downgrade level is assigned; when the number of replaceable core services equals the number of affected core services, a mild downgrade level is assigned; when the number of replaceable core services is greater than 0 and less than the number of affected core services, a moderate downgrade level is assigned; when the number of replaceable core services is 0, a severe downgrade level is assigned.

[0078] It should be noted that core services refer to DPU services whose core service identifiers are pre-written in the business deployment configuration; the allowable interruption range is determined by the difference between the number of affected core services and the number of replaceable core services.

[0079] Based on the downgrade level, disconnect the access path between the affected DPU service and the production business, lock the corresponding service hosting resources, and retain the running path of the unaffected service.

[0080] Furthermore, according to the degradation level corresponding to each node to be downgraded, the access ports, forwarding rules, service binding relationships, and session connections between the affected DPU service and the production business are retrieved; the access range that needs to be disconnected is determined based on the degradation level, a shutdown command is issued to the corresponding access port, the forwarding entry corresponding to the affected DPU service is deleted, the service binding between the production business and the affected DPU service is released, and the session connection that still points to the affected DPU service is terminated.

[0081] Based on the resource identifiers corresponding to the affected DPU services, locate the processing, storage, and communication resources that carry the affected DPU services, write isolation identifiers for the corresponding resources, suspend new business requests, and prevent the corresponding resources from being reallocated by other production services; select service paths from the shared DPU dependency graph that have not passed through the faulty components and whose probes respond normally, retain the corresponding access ports, forwarding rules, service binding relationships, and session connections, so that the unaffected DPU services can continue to carry production services.

[0082] A rescue activation command is sent to the node to be downgraded via an independent escape route, controlling the node to load the pre-set rescue operation environment and switch to the rescue operation state corresponding to the downgrade level.

[0083] Furthermore, a rescue initiation command is generated based on the identifier of the node to be downgraded, the downgrade level, and the rescue operation environment identifier, and sent to the corresponding node to be downgraded via an out-of-band management link or an independent management network; the node to be downgraded verifies the source of the command and the node identifier, and returns a command confirmation message upon successful reception; the node to be downgraded stops sending new business requests to the affected DPU service, saves the current necessary operating state, and reads the corresponding rescue operation environment from the local preset partition or independent startup storage.

[0084] Load the rescue operation environment into an independent operating space, switch the startup entry point, and start the basic operation services, network services, and business takeover services in the rescue operation environment; retain the business functions that are allowed to continue running according to the downgrade level, close the interfaces and tasks that depend on the affected DPU services, and transfer business requests to the rescue operation environment; when the basic services have started and the business takeover services can respond to requests, mark the nodes to be downgraded to the corresponding rescue operation status, and return rescue startup completion information carrying the node identifier and rescue operation status; verify the operation status of each node to be downgraded according to the rescue startup completion information, and trigger the repair of faulty components after all nodes to be downgraded have entered the rescue operation status.

[0085] After the node to be downgraded enters the rescue operation state, the corresponding repair strategy for the faulty component is adopted to stop the faulty component from continuing to carry production requests, clear the abnormal operation state, and perform component recovery operation to restore the faulty component.

[0086] Furthermore, after confirming that each node to be downgraded has entered the rescue operation state, the corresponding repair strategy is retrieved according to the fault component identifier; the access entry for the fault component to receive production requests is closed, the queue of requests that have not yet been executed is cleared, and the session connection between the fault component and the production business is terminated to avoid continuing to carry production requests during the repair period.

[0087] Read the running status, resource usage status and service configuration of the faulty component, release the abnormally occupied processing resources, storage resources and communication resources, delete the abnormal cache and failed sessions, and restore the service configuration to the most recent valid configuration; restart the faulty component, reload the service program or switch to the standby component according to the repair strategy, and mark the faulty component as recovered after the component has started up and its running status has returned to normal.

[0088] Send a management probe request to the recovered faulty component, verify the operating conditions of the faulty component based on the probe response, and mark the faulty component as recoverable when the faulty component meets the DPU service loading requirements.

[0089] Furthermore, it sends management probe requests carrying component identifiers and probe types to the recovered faulty components, receives the running status, resource usage status, service port status, management link status, and most recent startup information returned by the faulty components, aggregates various probe responses according to the faulty component identifier, and checks whether the probe responses are complete and whether they are returned within a preset time.

[0090] The running status, available resources, service ports, and management links are compared with the DPU service loading conditions one by one. When the faulty component is in normal operation, the remaining resources meet the service loading requirements, the service port can establish a connection, the management link remains connected, and no new abnormal records are detected, the faulty component is written into the list of recoverable components and marked as recoverable. If any condition is not met, the isolation status is retained and the unmet loading conditions are recorded.

[0091] It should be noted that the loading conditions for the DPU service are set according to the processing resources, storage resources, port connections, and management links required for the normal startup of the DPU service.

[0092] After a faulty component is marked as recoverable, a candidate recovery environment is created in a resource area isolated from production operations. The recovered DPU service and its running configuration are loaded into the candidate recovery environment, the dependent connections corresponding to the DPU service are restored, and the production access entry is closed. When the DPU service is started and can receive verification requests, a shared DPU candidate recovery state is formed.

[0093] Furthermore, after a faulty component is marked as recoverable, processing resources, storage space, and communication ports are allocated from a resource area isolated from production operations, and an independent running space is created according to the recovered DPU service identifier; the corresponding DPU service program and the most recent valid running configuration are read from the service image library and configuration storage area, and the DPU service program and running configuration are loaded into the independent running space.

[0094] Based on the service relationships recorded in the shared DPU dependency graph, re-establish the dependency connections between the DPU service and the faulty component, associated DPU services, and service-hosting resources, and confine each dependency connection to the isolated resource area; keep the production access port closed and prohibit production requests from entering the candidate recovery environment; start the recovered DPU service and check the service process, dependency connections, and verification port in sequence; when the DPU service has started, the dependency connections are normal, and the verification port can receive requests, record the DPU service identifier, running configuration, dependency connections, and startup status to form the shared DPU candidate recovery state.

[0095] S4. Based on the shared DPU candidate recovery status, when the replay response of the DPU service meets the normal response consistency condition and the completion status of the boundary request is determined, a valid service recovery token bound to the DPU service is generated, and the computing node is controlled to restore the shared DPU service path according to the service recovery token, thereby obtaining the shared DPU fault degradation switching information.

[0096] Retrieve loaded DPU services from the shared DPU candidate recovery status, select the service requests corresponding to the normal operation period, and replay the service requests to the candidate recovery environment according to the original call relationship; align the replay response with the corresponding normal response, compare the response characteristics, and write a replay consistency flag for the corresponding DPU service when the response difference is within the preset consistency range.

[0097] Furthermore, the system reads the identifiers, runtime configurations, and dependent connections of the started DPU services from the shared DPU candidate recovery status. It then filters the corresponding service requests from the service interaction records during normal operation according to the DPU service identifier. Complete service requests with request types, request parameters, and dependent objects are retained and arranged according to the original request time and call order to form a replay request sequence for each DPU service. Based on the caller, callee, and call order in the service interaction records, the replay requests are sent sequentially to the candidate recovery environment. After the previous service request returns a response, the data that needs to be transmitted in the response is written into the next service request according to the original call relationship. The system also records the request identifier, response content, response status, completion status, and response duration for each replay request.

[0098] Align the replay response with the normal response during normal operation according to the request identifier, and compare the response status, key response fields, request completion status and response duration respectively. When the response status and completion status are consistent, the differences in key response fields and the deviation in response duration are all within the same range, write a replay consistency flag for the corresponding DPU service. If any response feature exceeds the consistency range, record the corresponding difference item and do not write a replay consistency flag.

[0099] The consistency deviation of the DPU service replay response is calculated using the following formula:

[0100] ;

[0101] in, Indicates DPU service The smaller the value of the playback response consistency deviation, the more consistent the playback response is with the normal response; This indicates the DPU service that is accepted for replay verification in the candidate recovery environment; Indicates the sequence number of the playback request; Indicates DPU service The corresponding total number of replay requests; Indicates the sequence number of the key response field; Indicates the total number of numeric key response fields that participated in the consistency verification; Indicates the first The first in the replay response Key response field values; This indicates that the request identifier is used to retrieve the service interaction record during normal operation. The normal response corresponding to the replay request, and the first replay request extracted from the normal response. Baseline values ​​for key response fields; Indicates the first The allowable difference range of each key response field is determined by reading the field allowable error from the DPU service interface protocol and calculating the 95th percentile of the absolute deviation of the values ​​of similar response fields relative to the baseline value during normal operation, and taking the larger of the two values. Indicates the first The actual response time of a replay request in the candidate recovery environment; Indicates the period from normal operation to the first The baseline response time is extracted from the records of similar requests corresponding to each replay request. If there are multiple records, the average or median is taken. This indicates the allowable deviation in response time, which is set based on the latency fluctuation range of similar requests during normal operation and the allowable latency of the business.

[0102] It should be noted that the consistency range is set based on the response fluctuation range of similar service requests during normal operation, the allowable error of fields, and the allowable deviation of response time.

[0103] Filter out boundary requests that did not return explicit completion information from the DPU service interaction records before and after the failover, and track the corresponding execution records and status changes based on the request identifier; when the execution record and status change correspond to each other, mark the boundary request as completed; when no execution trace is found and resubmission is allowed, mark the boundary request as incomplete; when there are still conflicts in the completion status, retain the uncertain mark.

[0104] Furthermore, from the DPU service interaction records before and after the failover, requests that were sent to the DPU service but did not return success, failure, or cancellation information are filtered out and sorted according to the request identifier, DPU service identifier, and sending time to obtain the boundary requests to be verified. For each boundary request, the DPU service execution log, data write record, resource change record, and business status record are queried according to the request identifier, and the execution action and status change are compared in chronological order. When the execution log records that the request has been executed and the corresponding data, resource, or business status has changed in a way consistent with the request content, the boundary request is marked as completed.

[0105] If no request execution log, data write record, or status change is found, verify whether the boundary request has an idempotent flag and whether the corresponding DPU service allows resubmission. If the resubmission condition is met, mark the boundary request as incomplete. If the execution log shows that the request has been executed, but no corresponding status change is found, or the status has changed but execution record is missing, retain the uncertain flag and prohibit the boundary request from being directly resubmitted.

[0106] It should be noted that a boundary request refers to a service request that has been sent during failover but whose completion status has not been explicitly confirmed; the conditions for repeated submission are set based on whether the request is idempotent, whether repeated execution will cause business status conflicts, and the request processing rules of the DPU service.

[0107] When the DPU service has a replay consistency flag and all corresponding boundary requests have obtained completed or incomplete flags, the recovery verification information corresponding to the DPU service is collected, the recovery verification information is bound to the DPU service, and a valid service recovery token is generated. When the replay response does not meet the normal response consistency conditions or there are boundary requests with uncertain flags, the DPU service is kept in the candidate recovery environment and no valid service recovery token is generated.

[0108] Furthermore, the replay consistency flag and boundary request status corresponding to each DPU service are read, and the boundary requests are verified one by one to see if they have all been written with the completed or incomplete flag. When the DPU service has a replay consistency flag and there are no boundary requests with uncertain flags, the DPU service identifier, runtime configuration version, replay request identifier, replay response comparison information, boundary request identifier and corresponding completion status are extracted to form recovery verification information.

[0109] Write the recovery verification information into the recovery record of the corresponding DPU service and establish a binding relationship according to the DPU service identifier; write the token identifier, generation time, validity status and applicable service path into the bound recovery record to generate a valid service recovery token for the corresponding DPU service; when the DPU service fails to obtain the replay consistency mark, or any boundary request still has an uncertain mark, record the unmet verification conditions, keep the production access entry of the DPU service closed, and continue to keep the DPU service in the candidate recovery environment without generating a valid service recovery token.

[0110] It should be noted that a valid service recovery token is used to indicate that the corresponding DPU service has passed the response consistency verification, and that the completion status of each boundary request before and after the failover is clear.

[0111] Identify the corresponding DPU service from the service recovery token, determine the compute nodes that depend on the corresponding DPU service based on the shared DPU dependency graph, and verify whether the service recovery tokens that the compute nodes depend on are all in a valid state.

[0112] Furthermore, the DPU service identifier and token status are read from each service recovery token, and the corresponding service node in the shared DPU dependency graph is located according to the DPU service identifier. The dependency edges from the service node to the compute node are searched level by level to extract the compute node that calls the corresponding DPU service, and the dependency relationship between each compute node and the DPU service is recorded.

[0113] For each compute node, all DPU services that the compute node depends on for normal operation are aggregated, and the corresponding service recovery token is found according to the DPU service identifier. The existence of the service recovery token, the validity of the token status, and the consistency of the DPU service bound to the token are checked one by one. When all DPU services that the compute node depends on have corresponding valid service recovery tokens, the compute node is marked as meeting the service recovery conditions. If there is a missing service recovery token, an invalid token status, or a token that does not match the DPU service, the compute node will continue to be kept in the rescue operation state.

[0114] When the recovery conditions are met, the locking of the corresponding service-bearing resources and the isolation of the service path are released, and the access entry of the computing node is switched to the recovered shared DPU service path.

[0115] Furthermore, when the computing node meets the service recovery conditions, based on the DPU service identifier and applicable service path in the valid service recovery token, the corresponding service bearer resource, access port, forwarding rules, and service binding relationship are found; the lock identifier of the service bearer resource is cleared, the allocation permission of the corresponding resource is restored, the access port and forwarding rules corresponding to the shared DPU service after recovery are re-enabled, and the service binding between the computing node and the DPU service is established; the computing node is stopped from forwarding new business requests to the rescue operation environment, the access entry of the computing node is pointed to the shared DPU service path after recovery, and requests that have entered the rescue operation environment continue to be completed.

[0116] When service access is normal, the control computing node exits the rescue operation state. When service access is abnormal, the corresponding service recovery token is revoked, and the computing node is switched back to the rescue operation environment. The fault location, node degradation and service recovery information are collected to obtain the shared DPU fault degradation and switching information.

[0117] Furthermore, after switching the access entry point of the computing node to the restored shared DPU service path, a preset access request is sent to the corresponding DPU service, and it is checked whether the request is returned within the specified time, whether the response status is normal, and whether the key response content is complete. When the service access is normal, the rescue operation environment is stopped from receiving new business requests, and the received requests are waited for to be executed. The rescue operation entry point is then closed, and the computing node is marked as being in normal operation.

[0118] When service access times out, response status is abnormal, or critical response content is missing, the validity status of the corresponding service recovery token is changed to invalid, the access entry point of the recovered shared DPU service is closed, and the rescue operation environment is re-enabled; business requests of the compute node are transferred to the rescue operation environment, and the rescue operation status of the corresponding compute node is restored; the faulty component, abnormal path, affected DPU service, node to be downgraded, downgrade level, service recovery token status, and compute node switching status are collected and written into the same record according to the time of failure and the processing order to obtain the shared DPU fault degradation and switching information.

[0119] This embodiment also provides a computer device applicable to the fault degradation switching method based on a shared DPU, comprising: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the fault degradation switching method based on a shared DPU as proposed in the above embodiment.

[0120] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.

[0121] This embodiment also provides a storage medium storing a computer program that, when executed by a processor, implements the fault degradation switching method based on a shared DPU as proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0122] In summary, this invention constructs a dependency graph based on shared DPU runtime association data, sends differential probes to different computing nodes and DPU services, and traces back multiple dependency paths corresponding to abnormal responses step by step to the common component where the first independent deviation occurred. This enables the association and location of faulty components, affected DPU services, and nodes to be degraded, providing a unified basis for subsequent resource isolation, node degradation, and service recovery. It also avoids including non-faulty services or irrelevant computing nodes in the handling scope, improving the accuracy and continuity of shared DPU fault handling in industrial cloud computing environments.

[0123] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A fault degradation switching method based on a shared DPU, characterized in that, include: Collect the shared DPU's operating status, node access status, and out-of-band management status, and record the DPU service interaction data during normal operation to form shared DPU operating association data; Construct a shared DPU dependency graph based on shared DPU runtime correlation data, send differential probes to different compute nodes and different DPU services, compare the probe responses with normal responses, identify the faulty components corresponding to multiple abnormal paths along the shared DPU dependency graph, and identify the affected DPU services and nodes to be degraded. Based on the faulty components, determine the degradation level of each node to be degraded, isolate the service carrying resources corresponding to the affected DPU services, and control the nodes to be degraded to load the rescue operation environment through an independent escape channel; after the nodes to be degraded enter the rescue operation state, repair the faulty components, and load the recovered DPU services into a candidate recovery environment isolated from production business to obtain a shared DPU candidate recovery state. Based on the shared DPU candidate recovery status, when the replay response of the DPU service meets the normal response consistency condition and the completion status of the boundary request is determined, a valid service recovery token bound to the DPU service is generated, and the computing node is controlled to restore the shared DPU service path according to the service recovery token, thus obtaining the shared DPU fault degradation switchover information.

2. The fault degradation and switching method based on a shared DPU as described in claim 1, characterized in that, The process of collecting shared DPU operating status, node access status, and out-of-band management status, and recording DPU service interaction data during normal operation to form shared DPU operating association data, is as follows: According to the collection cycle, acquire the shared DPU running status, node access status and out-of-band management status, and write the associated information for each status data; Retrieve DPU service interaction records during normal operation, and establish associations between service interaction records and corresponding status data based on the association information to form shared DPU operation association data.

3. The fault degradation and switching method based on a shared DPU as described in claim 2, characterized in that, The specific steps for constructing a shared DPU dependency graph based on shared DPU runtime association data and sending differential probes to different compute nodes and different DPU services are as follows: Extract the calling and carrying relationships between objects from the shared DPU runtime association data, and construct a shared DPU dependency graph with objects as graph nodes and object relationships as graph edges. Select compute nodes and DPU services as probe targets along the shared DPU dependency graph, send differential probe requests according to different probe paths, and record the corresponding probe responses.

4. The fault degradation and switching method based on a shared DPU as described in claim 3, characterized in that, The steps for comparing probe responses with normal responses, identifying the faulty components corresponding to multiple abnormal paths along the shared DPU dependency graph, and determining the affected DPU services and nodes to be degraded are as follows: Align each probe response with its corresponding normal response, compare response characteristics, retain probe responses that deviate from the normal response range, and mark the dependency links between the probe sending position and the abnormal response position along the shared DPU dependency graph to form an abnormal path; The system aggregates abnormal paths corresponding to different compute nodes and DPU services, traces back level by level along the dependency direction, and filters out common components that appear repeatedly in multiple abnormal paths and are located upstream of abnormal forks. It compares the response deviations of common components with those of upstream components. When the upstream component responds normally and the common component is the first to show a response deviation, the common component is identified as the faulty component. When the upstream component also shows a response deviation, the system continues to trace back upstream of the dependency until the faulty component that first showed an independent response deviation is identified. Starting from the faulty component, expand along the service dependency direction, extract the service path containing the faulty component, identify the DPU services with abnormal responses in the service path as the affected DPU services, and identify the compute nodes that depend on the affected DPU services and cannot switch to the normal service path as nodes to be degraded.

5. The fault degradation and switching method based on a shared DPU as described in claim 4, characterized in that, The steps are as follows: Based on the faulty components, determine the degradation level of each node to be degraded, isolate the service-bearing resources corresponding to the affected DPU services, and control the nodes to be degraded to load the rescue runtime environment through an independent escape channel. Based on the abnormal path corresponding to the faulty component, determine the fault impact range and service replacement capability of each node to be downgraded, and match the corresponding downgrade level according to the downgrade rules; Disconnect the access path between the affected DPU service and the production business according to the downgrade level, lock the corresponding service carrying resources, and retain the running path of the unaffected service; A rescue activation command is sent to the node to be downgraded via an independent escape route, controlling the node to load the pre-set rescue operation environment and switch to the rescue operation state corresponding to the downgrade level.

6. The fault degradation and switching method based on a shared DPU as described in claim 5, characterized in that, After the node to be degraded enters the rescue operation state, the faulty components are repaired, and the recovered DPU service is loaded into the candidate recovery environment isolated from the production business to obtain a shared DPU candidate recovery state. The specific steps are as follows: After the node to be downgraded enters the rescue operation state, the repair strategy corresponding to the faulty component is adopted to stop the faulty component from continuing to carry production requests, clear the abnormal operation state, and perform component recovery operation to restore the faulty component. Send a management probe request to the recovered faulty component, verify the operating conditions of the faulty component based on the probe response, and mark the faulty component as recoverable when the faulty component meets the DPU service loading requirements; After a faulty component is marked as recoverable, a candidate recovery environment is created in a resource area isolated from production operations. The recovered DPU service and its running configuration are loaded into the candidate recovery environment, the dependent connections corresponding to the DPU service are restored, and the production access entry is closed. When the DPU service is started and can receive verification requests, a shared DPU candidate recovery state is formed.

7. The fault degradation and switching method based on a shared DPU as described in claim 6, characterized in that, Based on the shared DPU candidate recovery status, when the replay response of the DPU service meets the normal response consistency condition and the completion status of the boundary request is determined, a valid service recovery token bound to the DPU service is generated. The specific steps are as follows: Retrieve loaded DPU services from the shared DPU candidate recovery status, select the service requests corresponding to the normal operation period, and replay the service requests to the candidate recovery environment according to the original call relationship; Align the replay response with the corresponding normal response, compare the response characteristics, and write a replay consistency flag for the corresponding DPU service when the response difference is within the preset consistency range. Filter out boundary requests that did not return explicit completion information from the DPU service interaction records before and after the failover, and track the corresponding execution records and status changes based on the request identifier; when the execution record and status change correspond to each other, mark the boundary request as completed; when no execution trace is found and resubmission is allowed, mark the boundary request as incomplete; when there are still conflicts in the completion status, retain the uncertain mark. When the DPU service has a replay consistency flag and all corresponding boundary requests have obtained completed or incomplete flags, the recovery verification information corresponding to the DPU service is collected, the recovery verification information is bound to the DPU service, and a valid service recovery token is generated. If the replay response does not meet the normal response consistency conditions or there are boundary requests with uncertain markers, the DPU service remains in the candidate recovery environment and no valid service recovery token is generated.

8. The fault degradation and switching method based on a shared DPU as described in claim 7, characterized in that, The process of restoring the shared DPU service path by controlling the computing node based on the service recovery token, and obtaining shared DPU failure degradation and handover information, is as follows: Identify the corresponding DPU service from the service recovery token, determine the compute node that depends on the corresponding DPU service based on the shared DPU dependency graph, and verify whether the service recovery tokens that the compute node depends on are all in a valid state. When the recovery conditions are met, the locking of the corresponding service-bearing resources and the isolation of the service path are released, and the access entry of the computing node is switched to the recovered shared DPU service path; When service access is normal, the control computing node exits the rescue operation state. When service access is abnormal, the corresponding service recovery token is revoked, and the computing node is switched back to the rescue operation environment. The fault location, node degradation and service recovery information are collected to obtain the shared DPU fault degradation and switching information.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the fault degradation switching method based on any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the fault degradation switching method based on any one of claims 1 to 8.