Server fault intelligent diagnosis and logic self-healing system based on AI multi-modal feature fusion

The server fault intelligent diagnosis and logic self-healing system, which integrates AI multimodal features, solves the problem of secondary logic oscillation caused by local repairs in large-scale server clusters. It realizes intelligent risk diagnosis and self-healing under global topology constraints, ensures steady-state return of the system, and improves the operational continuity and stability of computing nodes.

CN121998622APending Publication Date: 2026-05-08SHANGHAI TEHUA COMPUTER SYST INTEGRATION CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI TEHUA COMPUTER SYST INTEGRATION CO LTD
Filing Date
2026-01-28
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing technologies cannot calculate the geometric deviation between the operating characteristics and the initial design manifold in real time in large-scale server clusters. This leads to local repair behaviors disrupting the system's steady state, causing secondary logical oscillations and network-wide topology collapse, and lacks the dynamic regression capability for global logical consistency.

Method used

A server fault intelligent diagnosis and logical self-healing system based on AI multimodal feature fusion is adopted. Through multimodal telemetry data acquisition, operation status benchmark model construction, operation deviation identification, logical self-healing engine and operation status consistency audit module, it realizes intelligent risk diagnosis and self-healing of server cluster under global topological constraints. It uses the principle of minimizing topological potential energy and Laplace operator to ensure that the system returns to the preset steady state.

Benefits of technology

It enables real-time risk identification and self-healing of server clusters, avoids secondary logic oscillations caused by partial repairs, ensures that the system deterministically returns to the preset steady state after being disturbed, and improves the operational continuity and stability of large-scale computing nodes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121998622A_ABST
    Figure CN121998622A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of server cluster operation and maintenance risk assessment, and discloses an AI multi-modal feature fusion-based server fault intelligent diagnosis and logic self-healing system, which comprises the following steps of: obtaining a multi-modal telemetry data stream containing a utilization rate, a memory allocation gradient and connection relation complexity; extracting an architecture design benchmark constraint to construct a running state multi-dimensional benchmark model; calculating a topology deviation value of a real-time state vector deviating from the model, decoupling steady-state deviation and dynamic disturbance by utilizing characteristic decomposition logic, and positioning a logic abnormal node; determining a compensation parameter set based on a consistency objective function minimization rule, and driving correlation node parameters to converge to design constraints; according to the method, node weights are coordinated through a global topology correction mechanism, secondary logic oscillation caused by local repair is avoided, and cluster operation continuity is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of server cluster operation and maintenance risk assessment technology, and in particular relates to a server fault intelligent diagnosis and logical self-healing system based on AI multimodal feature fusion. Background Technology

[0002] Current server cluster operation monitoring typically employs a perception architecture based on telemetry data and alarm logs. By acquiring processor utilization, memory resident status, and network communication status, and combining this with preset alarm thresholds, anomaly detection is performed to trigger corresponding early warning processes in the early stages of a risk. As the complexity of business logic increases, computing nodes exhibit strong logical coupling and temporal correlations, making a single fault signal often involve multi-dimensional risk characteristics. In large-scale dynamic cluster environments, traditional risk handling mechanisms mainly perform parameter calibration on local components by calling fixed repair scripts. However, there are highly complex topological dependencies within the cluster, and repair actions for local anomalies often generate unquantified logical stress. Because the global constraints of the business topology are ignored, such isolated repair behaviors can easily disrupt the existing resource balance of the system, causing nonlinear disturbances in the business links, which in turn can trigger secondary logical oscillations or even network-wide topology collapse.

[0003] Currently, the industry is attempting to improve the accuracy of self-healing by expanding expert experience rules or introducing more dimensional monitoring indicators. Analysis shows that these methods are essentially static pattern matching approaches, unable to calculate the geometric deviation between operational characteristics and the initial design manifold in real time. When faced with dynamic distortions in operational characteristics, existing solutions lack dynamic regression capabilities for global logical consistency, leading to uncertainty in the repair process. This contradiction between the effectiveness of local repair and global stability becomes the fundamental problem restricting the risk management efficiency of large-scale server clusters. For example, Chinese invention patent publication number CN112988444A... A method for diagnosing server cluster faults was developed. When automatic diagnosis fails, keyword matching is used to automate case reporting and work order distribution. The technology is essentially a known pattern matching process compensation method, focusing on closed-loop management of the diagnosis process. Non-fault root causes are dynamically decoupled at the global topology manifold level. However, it lacks the ability to analyze steady-state failures and dynamic disturbances caused by high-frequency business load switching. It cannot use real-time operating characteristics and initial design manifold geometric deviations to drive parameter convergence. The static rule diagnosis logic lacks global logical consistency and dynamic regression capability. The self-healing decision makes it difficult to ensure that the system state returns to the preset logical steady state after being disturbed.

[0004] Therefore, how to achieve intelligent risk diagnosis and self-healing instruction arrangement under global topological constraints, and ensure that the system state can deterministically return to the preset logical steady state after being disturbed, is the technical problem to be solved by this invention. Summary of the Invention

[0005] This invention provides a server fault intelligent diagnosis and logical self-healing system based on AI multimodal feature fusion, comprising: The multimodal telemetry data acquisition module is used to acquire multimodal telemetry data streams from the target server cluster. The multimodal telemetry data streams include processor utilization, memory allocation gradients, and network connectivity complexity. The runtime baseline model building module is used to perform semantic parsing on the system architecture configuration document to extract design baseline constraints, and map processor utilization, memory allocation gradient and network connection complexity into parameterized nodes in the logical topology, thereby constructing a multi-dimensional runtime baseline model that satisfies the design baseline constraints. The operation deviation identification module is used to extract the real-time state vector of the multimodal telemetry data stream, calculate the topological deviation value of the real-time state vector from the multidimensional benchmark model of the operation state, and extract the steady-state deviation component representing hardware failure and the dynamic disturbance component representing time instability by performing feature decomposition on the topological deviation value, so as to lock the logically abnormal nodes in the server cluster. The logic self-healing engine module is used to calculate a set of compensation parameters for logically abnormal nodes to make the global logical state tend to be stable based on the consistency objective function minimization rule. It also uses the compensation parameter set to adjust the memory allocation weight and request distribution step size of related nodes. Through parameter iteration convergence, the logical operation layer of the server cluster meets the design baseline constraints. The runtime consistency audit module is used to audit the feature spectrum distribution of the reconstructed logic graph using the Laplace operator, verify the consistency between the real-time feature spectrum and the original baseline feature spectrum, and thus confirm that the business logic link has been restored to the baseline runtime state.

[0006] Preferably, when constructing the multidimensional benchmark model of the running state, the benchmark model construction module defines the processor utilization as the flow dimension feature of the logical topology and the memory allocation gradient as the load dimension feature of the logical topology by performing feature extraction operations, and orthogonally processes the flow dimension feature and the load dimension feature using design benchmark constraints.

[0007] Preferably, when performing feature decomposition, the deviation identification module constructs a deviation covariance matrix, identifies the principal components with monotonically evolving characteristics in the topological deviation values ​​as steady-state deviation components, and defines the high-frequency components in the topological deviation values ​​as dynamic disturbance components.

[0008] Preferably, the logic self-healing engine module includes a strategy-driven submodule, which is used to retrieve repair instruction sequences from the fault handling knowledge graph according to the hierarchy of the logic anomaly node. The repair instruction sequences include logic parameter reset instructions, business process migration instructions, and device logic state change instructions.

[0009] Preferably, when performing feature spectrum distribution audit, the runtime consistency audit module determines the audit result by calculating the generalized Euclidean distance between the real-time feature spectrum and the original reference feature spectrum. The determination rule follows the formula below: ,in, For generalized Euclidean distance, The total number of feature dimensions. For the first Weight coefficients of dimensional features For the first in the real-time feature spectrum 1 eigenvalue, The first in the original benchmark feature spectrum 1 eigenvalue, This is the preset consistency judgment threshold.

[0010] Preferably, the logic self-healing engine module also includes an execution monitoring submodule. The execution monitoring submodule is used to trigger a secondary feature comparison operation after the repair instruction sequence is executed, and to calculate the reduction in global connection complexity of the server cluster after repair, so as to verify the disappearance of the topology deviation value.

[0011] Preferably, the multimodal telemetry data acquisition module is also used to acquire system log alarm information and perform text vectorization processing on the system log alarm information to serve as input variables for the operation deviation identification module to determine the root cause.

[0012] Preferably, when performing adjustment operations, the logic self-healing engine module coordinates the load weights of each parameterized node by introducing a parameter compensation operator into the second-order constraint equation to eliminate secondary logic oscillations caused by local parameter changes.

[0013] Preferably, the operating state baseline model construction module also includes a semantic extraction unit, which is used to perform path recognition on the system logic specification document using a text recognition algorithm to determine the logical boundaries of each parameterized node in the design baseline constraints.

[0014] Preferably, the runtime consistency audit module also includes a link verification unit. The link verification unit is used to detect the response latency of the business logic link by simulating business requests based on the results of the feature spectrum distribution audit. The unit of response latency is ms, so as to ensure that the data processing environment returns to the baseline runtime state.

[0015] Compared to existing technologies, the server fault intelligent diagnosis and logical self-healing system based on AI multimodal feature fusion of this invention has the following advantages: 1. In intelligent diagnosis of server faults, a global assessment mechanism for operational risks of computing environments is established by mapping multimodal telemetry data of server clusters to a logical steady-state manifold in a high-dimensional space. By calculating the topological geodesic distance between the real-time operational feature vector and the boundary of the healthy subspace, the risk of systematic deviations can be pre-identified. This risk measurement method based on spatial geometric relationships enables potential hidden risks caused by logical coupling between nodes to be objectively presented in the form of coordinate offsets, thereby overcoming the perception lag problem of traditional monitoring methods when facing large-scale dynamic clusters.

[0016] 2. The logic constraint solving engine, built on the principle of minimizing topological potential energy, transforms the generation process of self-healing instructions into solving the convergence path problem that minimizes global disturbances. By introducing parameter compensation operators into the second-order constraint equations, it coordinates and fine-tunes the load weights and memory residency gradients of associated nodes, enabling damaged logic links to collaboratively return to the design baseline through physical annealing. This avoids the risk of secondary logic oscillations caused by traditional local repair actions, ensuring that the overall steady-state balance of the data processing system is maintained while eliminating fault characteristics.

[0017] 3. By using the Laplace operator to audit the spectral domain symmetry of the reconstructed logic graph, this invention provides a technical means to verify whether the computing environment has been restored to the baseline safe state. By verifying the consistency between the real-time topology feature values ​​and the original design manifold in the spectral domain, the reliability of the business link is confirmed from the dimension of the integrity of the underlying architecture, improving the continuity of the operation of large-scale computing nodes and reducing the probability of the system crashing again due to incomplete repair or residual logic distortion. Attached Figure Description

[0018] Figure 1 This is a flowchart of the server fault diagnosis and logic self-healing execution process based on multimodal feature fusion of the present invention; Figure 2 This is a diagram of the dual-layer closed-loop control architecture of the physical entity layer and the intelligent logic self-healing layer of this invention. Detailed Implementation

[0019] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0020] It should be noted that all directional and positional terms used in this invention, such as: up, down, left, right, front, back, vertical, horizontal, inner, outer, top, bottom, transverse, longitudinal, center, etc., are only used to explain the relative positional relationship and connection between components in a specific state (as shown in the accompanying drawings). They are only for the convenience of describing this invention and do not require that this invention be constructed and operated in a specific orientation. Therefore, they should not be construed as limiting this invention. In addition, the descriptions of "first," "second," etc., in this invention are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated.

[0021] In the description of this invention, unless otherwise explicitly specified and limited, the terms installation, connection, and linking should be interpreted broadly. For example, they can refer to fixed connections, detachable connections, or integral connections; they can refer to mechanical connections; they can refer to direct connections or indirect connections through an intermediate medium; they can refer to the internal connection of two components. For those skilled in the art, the specific meaning of the above terms in this invention can be understood according to the specific circumstances.

[0022] In the description of this specification, references to the terms "an embodiment," "some embodiments," "illustrative embodiments," "examples," "specific examples," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example, and the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0023] This invention provides an intelligent server fault diagnosis and logic self-healing system based on AI multimodal feature fusion, comprising a multimodal telemetry data acquisition module, an operational status benchmark model construction module, an operational deviation identification module, a logic self-healing engine module, and an operational status consistency audit module. It is used to perform operational risk assessment of server clusters, extracting multidimensional operational features of the server cluster and mapping them to a topology space to achieve the location of logically abnormal nodes and automatic repair under global topology constraints. Addressing the technical problem that discrete metrics such as processor utilization, memory allocation gradient, and network connection complexity generated during dynamic operation of server clusters are difficult to directly characterize system-level risks, the multimodal telemetry data acquisition module acquires multimodal telemetry data streams of the target server cluster, including processor utilization, memory allocation gradient, and network connection complexity reflecting the strength of logical dependencies between nodes. The acquired real-time data is converted into a vector sequence with timestamp indexes, serving as the feature field input for subsequent logic deviation analysis. During nonlinear mapping, a radial basis operator is used to perform spatial transformation on processor utilization, memory allocation gradient, and network connection complexity. The real-time state vector is determined by calculating the radially symmetric distribution strength between the current metric point and a preset center point set. The projected coordinates in Hilbert space, where As a real-time state vector, the bandwidth parameter of the kernel function is adaptively corrected based on the sampling standard deviation of the real-time telemetry data to ensure that the state vector fully carries the business fluctuation information with nonlinear characteristics. This digital procedure based on spatial projection establishes the logical starting point for multimodal feature fusion.

[0024] The multimodal telemetry data acquisition module performs spatial transformation to determine the real-time state vector. Simultaneously, the system monitors processor utilization and memory allocation gradients within a sliding window. The sliding window size is specifically set to 50 consecutive sampling periods, with window sliding triggered every 10 sampling points to ensure an 80% overlap of sampling data between consecutive windows, thus smoothing out service glitches. Simultaneously, during the 2-second standby period after power-on initialization, the system performs 20 no-load excitation samplings on the physical link, calculates the standard deviation of the output as the background noise baseline, and sets it as the initial bias term used to determine feature offsets in subsequent radial basis function calculations. The standard deviation of the sampling sequence is also calculated. With scaling factor The product is set as the radial basis function bandwidth parameter. , scaling factor The value is taken from 2.5 to 3.5. The kernel function is used to calculate the projected coordinates of the feature points in Hilbert space, and the sampling standard deviation is used. Calculation cycle and telemetry data sampling cycle Synchronization reflects the nonlinear characteristic shift caused by fluctuations in business load. Given the current engineering situation in large-scale cluster environments where the lack of dynamic health benchmarks makes it impossible to quantify operational characteristic distortions, the operational state benchmark model construction module performs semantic parsing on system architecture configuration documents to extract design benchmark constraints. This module maps processor utilization, memory allocation gradient, and network connectivity complexity to parameterized nodes in the logical topology, thereby constructing a multi-dimensional operational state benchmark model that satisfies the design benchmark constraints. In the specific construction process, feature extraction operations are performed, defining processor utilization as the traffic dimension feature of the logical topology and memory allocation gradient as the load dimension feature. The design benchmark constraints are then used to orthogonalize the traffic and load dimension features to establish the logical steady-state manifold when the system is in a healthy state. .

[0025] The operational status baseline model construction module identifies the interconnection description fields in the system architecture configuration document and transforms unstructured physical parameters into matrix weights through a preset digital conversion mapping table. Specifically, the values ​​in the physical link bandwidth field are linearly scaled according to a weight ratio of 0.2 for 1Gbps and 1.0 for 10Gbps. The number of requests per unit time in the logical call frequency field is taken as the base-10 logarithm and mapped to a weight range of 0.1 to 0.9 as the adjacency matrix. The initial weights of the non-zero elements are used to determine the degree matrix based on the static topology of the physical links. Using the formula The benchmark Laplace matrix is ​​calculated to characterize the topological structure of the multidimensional benchmark model in operation, where... As the baseline Laplace matrix, and These are a diagonal matrix and an adjacency matrix describing the strength of logical dependencies between nodes, respectively. The diagonal elements are the sum of the connection weights of the corresponding nodes. Due to the real challenge that high-frequency temporal instability in server clusters often masks the true hardware failure characteristics and leads to ambiguity in root cause determination, the operation deviation identification module extracts the real-time state vector of the multimodal telemetry data stream, calculates the topological deviation value of the real-time state vector from the multidimensional benchmark model of the operation state, and extracts the steady-state deviation component representing hardware failure and the dynamic disturbance component representing temporal instability by performing feature decomposition on the topological deviation value, thereby locking the logically abnormal nodes in the server cluster. When performing feature decomposition, the module constructs a deviation covariance matrix, identifies the principal components with monotonically evolving characteristics in the topological deviation value as steady-state deviation components, and defines the high-frequency components in the topological deviation value as dynamic disturbance components, thereby achieving decoupling analysis of steady-state failure and transient fluctuations.

[0026] Addressing the technical challenge that repairing local anomalies can easily disrupt global resource balance and trigger secondary logical oscillations, the logic self-healing engine module, based on the consistency objective function minimization rule, calculates a set of compensation parameters for logically abnormal nodes to bring the global logical state to a steady state. This set of compensation parameters is used to adjust the memory allocation weights and request distribution step sizes of associated nodes. Through parameter iteration and convergence, the logical execution layer of the server cluster satisfies design baseline constraints. The logic self-healing engine module includes a strategy-driven submodule that retrieves repair instruction sequences from the fault handling knowledge graph based on the hierarchy of logically abnormal nodes. These repair instruction sequences include logical parameter reset instructions, business process migration instructions, and device logical state change instructions. When performing adjustment operations, this module introduces parameter compensation operators into the second-order constraint equations to coordinate the load weights of parameterized nodes, absorbing the topological stress caused by local parameter changes. The strategy-driven submodule employs semantic matching rules based on topological geodesic distance. The system uses locked logically abnormal nodes as query seeds, retrieving repair instruction paths from the knowledge graph that are closest in physical distance and logical call frequency, and introducing structured compensation operators into the second-order constraint equations. ,in, The damping correction matrix is ​​derived from the current topological potential. The partial derivatives with respect to the node weights are used to offset the topological stress caused by process migration in real time within the parameter update step size, ensuring that the convergence path of the overall state energy does not cross the preset logical collapse threshold.

[0027] When the logic self-healing engine module drives the convergence of associated node parameters, it introduces a damping characteristic structured compensation operator into the second-order constraint equation. To mitigate topological stress, the requested distribution step size is mapped at the hardware control layer to the distribution weight step pulse width in the load balancer control message. Each unit change in step size corresponds to a 2% offset adjustment in the load balancer's distribution ratio. The current global topological potential energy is then calculated. The partial derivatives of the weights of the disturbed nodes are used to form a gradient array, which is then compared with the decay factor. The product is the increment for compensation of the iteration parameters, and the decay factor. Set within the range of 0.5 to 0.8, the system updates the memory allocation weights within the iteration step size until the global topological potential generated between two adjacent iterations is reached. When the absolute value of the difference is less than the convergence threshold of 0.001, updates stop, driving the server cluster's logical operation layer to regress to the design baseline. Due to the auditing flaw that conventional operations and maintenance rely solely on the disappearance of error signals as a sign of successful repair and cannot verify the integrity of the underlying architecture, the runtime consistency auditing module uses the Laplace operator to audit the feature spectrum distribution of the reconstructed logical graph. This verifies the consistency between the real-time feature spectrum and the original baseline feature spectrum. The audit result is determined by calculating the generalized Euclidean distance between the real-time feature spectrum and the original baseline feature spectrum, and the determination rule follows the following formula: ,in, For generalized Euclidean distance, The total number of feature dimensions. For the first Weight coefficients of dimensional features For the first in the real-time feature spectrum 1 eigenvalue, The first in the original benchmark feature spectrum 1 eigenvalue, To set a preset consistency threshold, this module also includes a link verification unit. Based on the audit results of the feature spectrum distribution, it probes the response latency of the business logic link by simulating business requests. The response latency is based on... Using the unit of measurement to confirm that the data processing environment has returned to the baseline operating state, the system confirms the reliability of the business logic link from the dimension of consistency of topological characteristic values ​​by executing the above audit procedures, thereby achieving closed-loop verification of the fault handling effect.

[0028] Example 1: In a scenario where a distributed high-performance computing cluster with 500 physical nodes is performing a real-time large-scale parallel data migration task, the processor efficiency of some nodes in the server cluster may randomly decline due to operational status. Upon detecting increased processor utilization, the operation and maintenance monitoring system performs process migration. This operation disrupts the balanced distribution of global resources, increases network connectivity complexity, and causes logical oscillations, creating a risk of topology collapse in the data flow path. Furthermore, before collecting telemetry data streams containing business topology information, the system performs data anonymization to eliminate privacy data associated with specific user accounts. For the aforementioned uncertain multi-dimensional failure scenario, the multi-modal telemetry data acquisition module collects the multi-modal telemetry data streams of the computing cluster in real time, extracting the processor utilization, memory allocation gradient, and network connectivity complexity of each node. The operational status baseline model construction module, referencing the aforementioned procedures, uses design baseline constraints stored in the system architecture configuration document to determine the current logical steady-state manifold of the cluster within the multi-dimensional feature space. By calculating the deviation of the current state vector of each node The topology deviation value is used to identify that the fault is caused by the distortion of the logic feature field due to the processor efficiency fluctuation. The fault location is changed from finding physical bad points to correcting the deviation of the overall logic manifold of the system.

[0029] After receiving the steady-state deviation component output by the operational deviation identification module, the logic self-healing engine module starts the logic constraint solving engine to execute instruction orchestration. This engine calculates the global topological potential energy while maintaining the global topological weight balance. Minimize the set of compensation parameters, where the global topological potential is... The calculation formula is as follows: ,in, This represents the global topological potential. For nodes With nodes The coupling weights between them For real-time logical distance, To establish the baseline distance, the real-time logical distance is calculated by dividing the real-time sampled value of processor utilization by 100% to obtain the dimensionless traffic component, dividing the current periodic value of memory allocation gradient by the maximum total capacity to obtain the load component, and dividing the current number of active connections by the designed maximum number of connections (65535) to obtain the topology density component. These three components are then weighted and summed in proportions of 0.5, 0.3, and 0.2. This value is updated every 10ms and stored in a circular buffer of length 5. The physical deviation of the current logical link is determined by calculating the moving average of the data in the buffer. Based on this compensation parameter set, the system issues correction instructions for the memory allocation weights of associated nodes and simultaneously adjusts the request distribution step size, causing the damaged service links to converge towards the design baseline point. This coordinated parameter adjustment absorbs the logical stress caused by local parameter changes, resolving the technical contradiction of global oscillations caused by local repairs. After the parameter iterative convergence process is completed, the runtime consistency audit module uses the Laplace operator to audit the feature spectrum distribution of the reconstructed logical graph, calculating the generalized Euclidean distance between the real-time feature spectrum and the original baseline feature spectrum. To assess the quality of repair, the generalized Euclidean distance The calculation formula is as follows: ,in, For generalized Euclidean distance, The total number of feature dimensions. For the first Weight coefficients of dimensional features For the first in the real-time feature spectrum 1 eigenvalue, The first in the original benchmark feature spectrum A characteristic value is used when the audit result meets the judgment criteria. At that time, among them The preset consistency threshold indicates that the business logic link of the computing cluster has recovered to a steady-state operating state that meets the design baseline constraints.

[0030] Example 2: In a server cluster operating environment where timing instability and telemetry data distortion are caused by high-frequency service load switching, this experiment addresses the technical problem of distinguishing between rigid deviations caused by hardware physical failures and elastic disturbances caused by service fluctuations. It employs a parallel computing simulation platform built using the Monte Carlo simulation method. Gaussian white noise with a signal-to-noise ratio of 25dB and 50Hz power frequency interference harmonics are superimposed onto the original operating metrics such as processor utilization, memory allocation gradient, and network connection complexity to simulate the data acquisition process under industrial electromagnetic conditions. The sampling period is... The determination follows the frequency bandwidth matching rule. When the characteristic field frequency bandwidth of the telemetry data is in the range of 0.1Hz to 10Hz, in order to meet the Nyquist sampling criterion and reserve 20% computational redundancy, the sampling period is... The time was set to 10ms to suppress signal aliasing while capturing transient features. The experimental group adopted the technical solution disclosed in this invention, control group 1 removed the feature decomposition logic in the operation deviation identification module, and control group 2 adopted the discrete trigger repair method. Both groups were applied to the real-time state vector containing noise interference. During the trial operation, the system monitored in real time whether the computing nodes deviated from the logical steady-state manifold. The path evolution trend was recorded, and the accuracy of each group in locating the logical abnormal node and the convergence time after executing the self-healing instruction were recorded. The experimental group used the deviation covariance matrix to decompose the real-time state vector into steady-state deviation components and dynamic disturbance components to filter out the interference of random noise on the root cause determination. The topological deviation value extracted by the control group 1 contained elastic distortion components, resulting in root cause location drift, as shown in Table 1.

[0031] Table 1: Comparison of Operation and Maintenance Performance Data for Different Test Groups

[0032] According to the data recorded in Table 1, the experimental group maintained a diagnostic accuracy of 98.24% when processing telemetry data streams containing 25dB noise, and the global topological potential was also stable. The convergence time was 342.56 ms, and the oscillation rate was 0.52%. By comparing the experimental group with the control group 1, the experimental group performed feature orthogonal decomposition to extract physically oriented steady-state deviation components, enabling the logic self-healing engine module to obtain a definite feature field input, thus improving the convergence efficiency by 33.15%. According to the formula... The defined global state energy constraint allowed the experimental group to avoid the secondary logic oscillations caused by discrete repair actions in control group 2; when the feature weight coefficients When the value is within the preset range of 0.1 to 0.9, the system performs orthogonalization processing; and when... When the value is set to 0.95 and exceeds the upper limit, the diagnostic accuracy drops to 85.67% and the convergence path exhibits nonlinear oscillations, confirming the rationality of the parameter distribution range. After the instruction execution is complete, the runtime consistency audit module calculates the feature spectrum of the reconstructed logic graph, and the measured generalized Euclidean distance is used. The value is 0.12, which is less than the preset consistency judgment threshold. Furthermore, the response latency of the business logic link returned to 12.84ms from 150.52ms, proving that the fault diagnosis and self-healing mechanism can achieve consistent calibration of the global architecture under noisy interference environment.

[0033] Example 3: This example combines Figures 1 to 2 This section describes a server fault intelligent diagnosis and logical self-healing system based on AI multimodal feature fusion, such as... Figure 1 As shown, the multimodal telemetry data acquisition module acquires real-time data streams containing utilization, memory allocation gradient, and connection complexity. Simultaneously, the runtime state benchmark model construction module extracts architecture design benchmark constraints to construct a multidimensional runtime state benchmark model. The generated real-time state vector and benchmark model constraints are input to the runtime deviation identification module, which calculates the topology deviation value and performs feature decomposition decoupling to locate logically abnormal nodes. The system passes the calculated topology deviation value to the logic self-healing engine module, which determines the compensation parameter set based on the consistency objective function minimization rule to drive the convergence of associated node parameters. Finally, the generated compensation parameter set and the adjusted state enter the runtime state consistency audit module, which uses the Laplace operator to perform spectral distribution audit to verify the consistency of the feature spectrum.

[0034] like Figure 2 As shown, the system architecture is divided into a physical operating environment entity layer and a logical self-healing computing environment intelligent layer. The target server cluster in the entity layer includes processors, memory, and network entities. This layer collects data through a multimodal telemetry data acquisition module and a system log alarm collection function (shown in the dashed box), and inputs the information unidirectionally to the intelligent layer through feature mapping and spatial transformation. The intelligent layer constructs a closed-loop control loop around the central fault handling knowledge graph. It includes an operating state benchmark model for establishing the logical steady-state manifold, an operating deviation identification module for performing feature decomposition and root cause locking, a logical self-healing engine for minimizing the consistency objective function, and a consistency audit module for performing feature spectrum verification. The output of the intelligent layer after processing is converted into topology parameter compensation instructions and fed back to the entity layer, thereby realizing closed-loop control and logical self-healing of the physical operating environment.

[0035] Example 4: In application scenarios where server clusters experience high-dimensional nonlinear oscillations due to network connectivity complexity crossing critical points, the real-time state vector... With logically steady manifold The topological deviation values ​​between the values ​​produce non-stationary random fluctuations. The deviation identification module executes a feature space reconstruction procedure based on singular value decomposition, which constructs the deviation covariance matrix. The calculation formula is as follows: ,in, The deviation covariance matrix, The total number of samples within the preset time window. For the first The real-time state vector at each moment. The reference state vector; the system's matrix Perform orthogonal decomposition by selecting the first The eigenvectors corresponding to the largest eigenvalues ​​are used to construct a projection matrix, which maps the topological deviation values ​​to a subspace defined by the steady-state deviation components. The value is determined by the cumulative contribution rate method, i.e., the current... The sum of the energy proportions of each characteristic value is not less than 85%, thereby filtering out dynamic disturbance components in the energy distribution at the tail; and a consistency judgment threshold is set for the load environment. For technical issues that are difficult to determine, the system initiates an adaptive parameter calibration procedure. During the quasi-steady-state period of the data processing environment, the operational consistency audit module continuously monitors the fluctuation amplitude of the real-time feature spectrum and calculates the statistical standard deviation of the feature spectrum sequence. Based on this statistical distribution characteristic, a threshold is determined. The calculation formula is as follows: ,in, The threshold for consistency determination This is the scaling factor, which is set to 3.0 under the conditions of this test. The statistical standard deviation of the feature spectrum sequence is used. By introducing calibration logic based on real-time background noise intensity, the system resets the judgment benchmark according to the new feature spectrum distribution when facing a step change in processor utilization, thus avoiding false triggering of the self-healing engine caused by load baseline drift.

[0036] When executing instruction orchestration, the logic self-healing engine module utilizes the logic constraint solving engine to solve for the global topological potential. Iterative optimization is performed, which decomposes the repair action into a sequence of parameter increments. In each iteration step, the system calculates the partial derivative of the current potential energy with respect to the weights of each node, and updates the memory allocation weights and request distribution step size along the negative gradient direction. The update procedure follows a linear approximation of the second-order constraint equation, that is, when the global topological potential energy... The iteration stops when the absolute value of the descent slope is less than 0.001. During the self-healing decision-making phase, the strategy-driven submodule retrieves causally related repair paths from the fault handling knowledge graph based on the topological coordinates of the logically abnormal nodes. When the memory allocation weight deviation of a memory-intensive node is found to exceed 15%, the system associates the corresponding business process migration instruction and smooths out the logical stress generated during the migration process by adjusting the request distribution step size. After completing the above parameter iteration, the runtime consistency audit module collects the reconstructed business topology data and calculates the reconstructed real-time feature spectrum using the Laplace operator. By performing generalized Euclidean distance calculation, the system obtains a final audit value of 0.15, which is lower than the dynamic threshold determined by the above calibration procedure. The dynamic threshold determined at this point is 0.42, indicating that the repaired logic graph and the original baseline manifold have re-achieved symmetry constraints in the spectral domain. The request-response latency measured synchronously by the link verification unit is stable at 15.0ms, and the latency jitter variance is reduced by 65.4%. This confirms that the business logic link of the computing cluster has eliminated the risk state caused by high-dimensional oscillations and returned to a steady-state equilibrium that satisfies the design baseline constraints.

[0037] Example 5: When the system faces the situation of first-time management of heterogeneous computing architecture or physical link reconstruction, the runtime state baseline model construction module executes the baseline calibration procedure. During the quiet period before service launch, it constructs synthetic telemetry flow excitation computing nodes with fixed load step size, and uses discrete sampling to extract the processor utilization, memory allocation gradient, and network connection complexity of each node in a zero-disturbance state. The extraction frequency is not less than 100Hz, and a baseline state vector is synthesized based on the average of 500 consecutive sets of samples. The system utilizes parameterized nodes to perform manifold learning operations in the feature space, and establishes a logically steady-state manifold by constructing a spectral distribution that satisfies design baseline constraints. The initial location is used to establish the reference coordinate system required for the operation of the diagnostic logic.

[0038] In application scenarios involving the filling of knowledge graph entries for fault handling and the weighting of repair paths, the logical self-healing engine module utilizes a controlled fault injection procedure to quantify the correlation strength between the abnormal feature field and the repair instruction sequence, and simulates the local topological potential energy caused by a single-point failure. The distortion is recorded in the second-order constraint equations to the response feedback of the parameter compensation operator to different repair commands. The command path that makes the absolute value of the slope of the overall situational energy decrease approach the 0.001 threshold is defined as the mapping entry of the corresponding root cause node, and the weight coefficient of each repair path is calculated accordingly. The system synchronizes the measured data to the policy-driven submodule with the measured data and the corresponding parameter step size, so that the instruction orchestration process has a logical judgment basis based on topology convergence performance.

[0039] Example 6: In a deployment procedure where a server cluster involves a specific topology architecture and requires the establishment of initial design constraints, the runtime baseline model construction module executes a semantic extraction procedure for configuration documents. It uses preset unstructured data parsing operators to identify interconnection description fields in the system architecture configuration documents, converting the physical link bandwidth and logical call frequency between server nodes into design baseline distances. ,in, To establish a baseline distance, the system extracts the forwarding latency and memory residency period of the Layer 3 switch from the design document. It then calculates the mutual exclusion weights of each node pair in the feature spectrum space using a linear weighting method, and uses this weighted average to construct the logically steady-state manifold. With structured input, when the number of physical nodes in the deployment environment expands from 500 to 1000, the system automatically injects new parameterized nodes into the logical topology by identifying the configuration descriptions of the newly added nodes, thus establishing the initial boundaries of the risk assessment process under different architectural scales.

[0040] In operational scenarios involving fault feature field injection and the need for dynamic correction of the fault handling knowledge graph, the logical self-healing engine module initiates a feedback entry update procedure. When the abnormal feature vector locked by the operational deviation identification module has no corresponding instruction path in the knowledge graph, the system utilizes the logical constraint solving engine to perform a global random search. By perturbing the memory allocation weights of associated nodes within a 5ms step, it probes the global topological potential. The gradient change, if the parameter adjustment action is detected to change the global topological potential energy If the parameter decreases monotonically over three consecutive sampling periods, the system defines the parameter adjustment sequence as a candidate self-healing path for the corresponding fault. The system then uses the feature spectrum distribution audit results after execution as confidence labels to store them in the fault handling knowledge graph. This procedure establishes a closed loop mapping from fault characteristics to verifiable instructions, enabling the self-healing logic of the server cluster to have a continuously evolving reference basis.

[0041] The embodiments of this application have been described above with reference to the accompanying drawings. Unless otherwise specified, the embodiments and features in the embodiments of this application can be combined with each other. This application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit of this application and the scope of protection of this invention, and all of these forms are within the protection scope of this application.

Claims

1. A server fault intelligent diagnosis and logical self-healing system based on AI multimodal feature fusion, characterized in that, include: The multimodal telemetry data acquisition module is used to acquire multimodal telemetry data streams from the target server cluster. The multimodal telemetry data streams include processor utilization, memory allocation gradients, and network connectivity complexity. The runtime baseline model building module is used to perform semantic parsing on the system architecture configuration document to extract design baseline constraints, and map processor utilization, memory allocation gradient and network connection complexity into parameterized nodes in the logical topology, thereby constructing a multi-dimensional runtime baseline model that satisfies the design baseline constraints. The operation deviation identification module is used to extract the real-time state vector of the multimodal telemetry data stream, calculate the topological deviation value of the real-time state vector from the multidimensional benchmark model of the operation state, and extract the steady-state deviation component representing hardware failure and the dynamic disturbance component representing time instability by performing feature decomposition on the topological deviation value, so as to lock the logically abnormal nodes in the server cluster. The logic self-healing engine module is used to calculate a set of compensation parameters for logically abnormal nodes to make the global logical state tend to be stable based on the consistency objective function minimization rule. It also uses the compensation parameter set to adjust the memory allocation weight and request distribution step size of related nodes. Through parameter iteration convergence, the logical operation layer of the server cluster meets the design baseline constraints. The runtime consistency audit module is used to audit the feature spectrum distribution of the reconstructed logic graph using the Laplace operator, verify the consistency between the real-time feature spectrum and the original baseline feature spectrum, and thus confirm that the business logic link has been restored to the baseline runtime state.

2. The server fault intelligent diagnosis and logical self-healing system based on AI multimodal feature fusion according to claim 1, characterized in that, When constructing a multidimensional benchmark model for the running state, the benchmark model construction module defines processor utilization as the flow dimension feature of the logical topology and memory allocation gradient as the load dimension feature of the logical topology by performing feature extraction operations. It also orthogonally processes the flow dimension feature and load dimension feature using design benchmark constraints.

3. The server fault intelligent diagnosis and logical self-healing system based on AI multimodal feature fusion according to claim 1, characterized in that, When performing eigenvalue decomposition, the deviation identification module constructs a deviation covariance matrix to identify principal components with monotonically evolving characteristics in the topological deviation values ​​as steady-state deviation components, and defines high-frequency components in the topological deviation values ​​as dynamic disturbance components.

4. The server fault intelligent diagnosis and logical self-healing system based on AI multimodal feature fusion according to claim 1, characterized in that, The logic self-healing engine module includes a strategy-driven submodule, which is used to retrieve repair instruction sequences from the fault handling knowledge graph according to the hierarchy of the logic anomaly node. The repair instruction sequences include logic parameter reset instructions, business process migration instructions, and device logic state change instructions.

5. The server fault intelligent diagnosis and logical self-healing system based on AI multimodal feature fusion according to claim 1, characterized in that, When performing feature spectrum distribution audit, the runtime consistency audit module determines the audit result by calculating the generalized Euclidean distance between the real-time feature spectrum and the original baseline feature spectrum. The determination rule follows the formula below: ,in, For generalized Euclidean distance, The total number of feature dimensions. For the first Weight coefficients of dimensional features For the first in the real-time feature spectrum 1 eigenvalue, The first in the original benchmark feature spectrum 1 eigenvalue, This is the preset consistency judgment threshold.

6. The server fault intelligent diagnosis and logical self-healing system based on AI multimodal feature fusion according to claim 4, characterized in that, The logic self-healing engine module also includes an execution monitoring submodule. After the repair instruction sequence is executed, the execution monitoring submodule triggers a secondary feature comparison operation and calculates the reduction in global connection complexity of the server cluster after repair, in order to verify the disappearance of the topology deviation value.

7. The server fault intelligent diagnosis and logical self-healing system based on AI multimodal feature fusion according to claim 1, characterized in that, The multimodal telemetry data acquisition module is also used to acquire system log alarm information and perform text vectorization processing on the system log alarm information to serve as input variables for the operational deviation identification module to determine the root cause.

8. The server fault intelligent diagnosis and logical self-healing system based on AI multimodal feature fusion according to claim 1, characterized in that, When performing adjustment operations, the logic self-healing engine module coordinates the load weights of each parameterized node by introducing parameter compensation operators into the second-order constraint equations to eliminate secondary logic oscillations caused by local parameter changes.

9. The server fault intelligent diagnosis and logical self-healing system based on AI multimodal feature fusion according to claim 1, characterized in that, The operational baseline model construction module also includes a semantic extraction unit, which uses text recognition algorithms to identify paths in the system logic specification document in order to determine the logical boundaries of each parameterized node in the design baseline constraints.

10. The server fault intelligent diagnosis and logical self-healing system based on AI multimodal feature fusion according to claim 1, characterized in that, The runtime consistency audit module also includes a link verification unit, which is used to detect the response latency of the business logic link by simulating business requests based on the audit results of the feature spectrum distribution.

Citation Information

Patent Citations

  • Processing method for fault diagnosis of server cluster

    CN112988444A