HBase client master-slave switching method and system based on fault perception

Through distributed fault monitoring and deep learning technology, combined with adaptive clustering and wavelet packet decomposition, the fault feature map of the HBase client is generated, and the misjudgment and misjudgment problems of the main and backup switching schemes in the existing technology are solved, efficient fault diagnosis and intelligent switching are achieved, and the reliability and availability of the system are improved.

CN119537484BActive Publication Date: 2025-05-16北京科杰科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510110298.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-16
Estimated Expiration
2045-01-23

AI Technical Summary

Technical Problem

The existing HBase client master-subsidiary switching scheme cannot effectively identify complex system abnormalities, which can easily cause misjudgment or misjudgment, be difficult to adapt to dynamically changing system environments, and be unable to discover potential risk points in time, resulting in unreliable services.

Method used

The distributed fault monitoring module collects multi-dimensional system index data, combines adaptive clustering and deep neural network to establish an abnormality detection model, performs data distribution learning and preprocessing, uses wavelet packet decomposition and principal component analysis to perform feature extraction, generates a system health assessment distribution matrix, combines density clustering and probability network to analyze the fault propagation path, generates a fault feature map, and intelligent switching is performed through deep strategy learning and double-buffer zero-copy mechanism.

Benefits of technology

It realizes accurate identification and timely warning of system abnormalities, improves the accuracy and real-time nature of fault detection, quickly identifys fault propagation paths and root causes, improves fault diagnosis efficiency and system availability, and reduces the impact of the switching process on business.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119537484B_ABST
    Figure CN119537484B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for HBase client active-standby switching based on fault perception, which relates to the technical field of database management, including: collecting system operation data through a distributed fault monitoring module, establishing an abnormality detection baseline model, performing fault diagnosis based on a system health assessment distribution matrix, generating a fault feature map, triggering switching according to risk level assessment results, and selecting the optimal standby node for sequential switching using a deep strategy learning algorithm, which can achieve fast and accurate fault detection, diagnosis and intelligent switching, and improve system reliability and performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of database management, and in particular to a method and system for HBase client master-slave switching based on fault perception. Background Art

[0002] With the rapid development of distributed database technology, HBase, as a highly reliable and high-performance distributed database system, has been widely used in large-scale data storage and processing scenarios. In practical applications, the HBase client is an important component that connects the application and the database server. Its stability and availability directly affect the operation of the entire system.

[0003] HBase clients usually use a master-slave architecture to ensure service continuity. When the master node fails, the client can switch to the standby node to continue providing services. Traditional master-slave switching solutions mainly rely on heartbeat detection and timeout mechanisms to determine the node status and perform failover through preset switching strategies.

[0004] The existing HBase client master-slave switching still has problems such as being unable to effectively identify complex system abnormalities, easily causing misjudgments or missed judgments, being difficult to adapt to dynamically changing system environments, and being difficult to accurately track the propagation of faults, and being unable to promptly discover potential risk points.

[0005] Therefore, a solution is urgently needed to solve the problems existing in the prior art. Summary of the invention

[0006] The embodiment of the present invention provides a method and system for HBase client master-slave switching based on fault perception, which can at least solve some problems existing in the prior art.

[0007] A first aspect of an embodiment of the present invention provides a method for HBase client master-slave switching based on fault perception, comprising:

[0008] Collect system operation data through the distributed fault monitoring module in the client and form a multi-dimensional system indicator data set through monitoring probes deployed in different network areas, perform data distribution learning on the multi-dimensional system indicator data set through an adaptive clustering algorithm and establish an anomaly detection baseline model, perform a sliding time window operation on the multi-dimensional system indicator data set based on the anomaly detection baseline model and the exponential moving average algorithm in combination with a low-pass filter to obtain pre-processed standard indicator data, extract time-frequency domain features from the standard indicator data through a multi-scale analysis method based on wavelet packet decomposition, perform dimensionality reduction in combination with a principal component analysis algorithm to obtain a multi-dimensional feature vector, and generate a system health assessment distribution matrix based on a pre-set deep neural network model and the anomaly detection baseline model;

[0009] Based on the system health assessment distribution matrix, the abnormal index is determined and compared with the adaptive threshold calculated by the density clustering algorithm. If it is greater than the adaptive threshold, the fault diagnosis model is triggered, and the fault-related data is collected in a targeted manner through the improved distributed tracing system. The fault-related data is constructed as a service call dependency graph based on the timing graph calculation algorithm combined with the gated graph attention mechanism. The abnormal propagation path is analyzed by a causal reasoning model based on a probability network in combination with the multi-dimensional feature vector, and a fault propagation chain is generated. The system log is obtained, and the system log and the fault propagation chain are associated with each other based on a deep semantic analysis model and a graph structure network. The context features are extracted and a fault feature map is generated in combination with the knowledge transfer technology. The fault feature map is multi-layered classified based on the integrated decision tree framework, the features of the abnormal propagation path are integrated, and the fault root cause analysis report and risk level assessment results are output in combination with the hierarchical generalization strategy. The switching strategy parameter set is generated in combination with the dynamic programming algorithm.

[0010] Based on the risk level assessment result, it is determined whether to trigger switching. If triggered, the intelligent switching execution module is started, and candidate nodes are screened from the global standby node pool through the consistent hashing algorithm combined with the virtual node mechanism and according to the switching policy parameter set. Based on the deep policy learning algorithm, the characteristic indicators in the fault root cause analysis report are used to construct a standby node scoring system. The decision process is optimized through the search tree and dynamic scoring is performed. The candidate node with the highest score value is used as the switching target node and the double buffer zero copy mechanism is initialized based on the fault propagation chain. The database operation request is prioritized in combination with the multi-level feedback queue algorithm, and sequential switching is performed based on the priority division result. The performance indicators in the switching process are compared with the system health assessment distribution matrix in combination with the timing prediction algorithm to model, and the switching policy parameter set is optimized until the switching is completed.

[0011] In an optional embodiment,

[0012] The distributed fault monitoring module in the client collects system operation data and forms a multi-dimensional system indicator data set through monitoring probes deployed in different network areas. The adaptive clustering algorithm is used to learn the data distribution of the multi-dimensional system indicator data set and establish an anomaly detection baseline model, including:

[0013] The distributed fault monitoring module collects system operation data, and the system operation data includes processor utilization index, memory occupancy index, disk performance index, and network performance index. The processor utilization index is obtained by collecting the user state utilization rate, system state utilization rate, input and output waiting rate, and idle rate of each processor core. The memory occupancy index is obtained by collecting the physical memory usage, virtual memory usage, cache usage, and buffer usage. The disk performance index is obtained by collecting the read and write rate, average response time, input and output queue length, and device utilization of each disk device. The network performance index is obtained by collecting the input and output bandwidth utilization, network delay, packet loss rate, and retransmission rate.

[0014] The distributed fault monitoring module collects the system operation data through monitoring probes deployed in different network areas to form a multi-dimensional system indicator data set, wherein the monitoring probes are composed of a first-level probe node, a second-level probe node, and a third-level probe node. The first-level probe node is configured according to the ratio of the number of racks and is responsible for local data collection. The second-level probe node is configured according to the number of computer rooms and is responsible for aggregating the data collected by the first-level probe node for preprocessing. The third-level probe node is configured after dividing the area according to the geographical location and network topology structure and is responsible for coordinating and managing the data collection tasks in the area. The first-level probe node, the second-level probe node, and the third-level probe node maintain connection through a heartbeat mechanism.

[0015] After the multidimensional system indicator data set is standardized and converted, the Euclidean distance between the sample and the cluster center is calculated based on the initial cluster center for classification. When the variance of the samples in the cluster exceeds the preset threshold, a new cluster center is added in the direction of the maximum variance in the current cluster. When the distance between adjacent cluster centers is less than the preset threshold, adjacent clusters are merged. If the moving distance of the cluster center is less than the preset convergence threshold for 100 consecutive iterations, the data distribution learning is considered to be completed.

[0016] Based on the results of data distribution learning, the mean and standard deviation of samples in each cluster are calculated, and the multiple range of the mean and the standard deviation is set as the normal value range of the current indicator. At the same time, the historical data is divided into hourly cycles to establish an hourly baseline, a daily baseline is established according to a daily cycle, and a monthly baseline is established according to a monthly cycle. The hourly baseline, the daily baseline, and the monthly baseline are weighted to form the anomaly detection baseline model.

[0017] In an optional embodiment,

[0018] Based on the anomaly detection baseline model and the exponential moving average algorithm, a sliding time window operation is performed on the multidimensional system indicator data set in combination with a low-pass filter to obtain preprocessed standard indicator data, time-frequency domain feature extraction is performed on the standard indicator data through a multi-scale analysis method based on wavelet packet decomposition, and a multi-dimensional feature vector is obtained by dimensionality reduction in combination with a principal component analysis algorithm, and a system health assessment distribution matrix is ​​generated based on a pre-set deep neural network model and the anomaly detection baseline model, including:

[0019] The distributed fault monitoring module performs a sliding time window operation on the multidimensional system indicator data set based on the anomaly detection baseline model and the exponential moving average algorithm in combination with a low-pass filter, fills the continuous indicator with linear interpolation of the two normal values ​​before and after, and fills the discrete indicator with the nearest neighbor normal value, distinguishes high-frequency change indicators from low-frequency change indicators based on the ratio of the standard deviation to the mean, and smoothes the high-frequency change indicators and the low-frequency change indicators with different smoothing factors to obtain pre-processed standard indicator data;

[0020] The time-frequency domain features of the standard indicator data are extracted by a multi-scale analysis method based on wavelet packet decomposition, the standard indicator data is decomposed into multiple frequency bands, the energy value of each frequency band is calculated, the frequency bands whose energy proportion is lower than a preset energy proportion threshold are eliminated, the mean, variance, skewness, kurtosis, maximum value and minimum value of the retained frequency bands are extracted as statistical features, and the energy ratio between different frequency bands is calculated based on the statistical features to obtain the time-frequency domain features;

[0021] The distributed fault monitoring module calculates the covariance matrix of the feature matrix based on the principal component analysis algorithm to obtain eigenvalues ​​and eigenvectors, selects eigenvectors corresponding to the eigenvalues ​​based on a preset variance contribution rate threshold to construct a conversion matrix, and maps the time-frequency domain features to the principal component space to obtain a multi-dimensional eigenvector;

[0022] The distributed fault monitoring module inputs the multi-dimensional feature vector into a preset deep neural network model, wherein the number of input layer nodes of the deep neural network model is the same as the number of principal components, the number of hidden layer nodes is set according to the ratio of the input dimensions, the number of output layer nodes is the same as the type of monitoring indicators, the parameters of the deep neural network model are trained based on normal state data, the multi-dimensional feature vector is input into the trained deep neural network model to obtain a health score, and normalized in combination with the anomaly detection baseline model to generate a system health assessment distribution matrix.

[0023] In an optional embodiment,

[0024] Based on the system health assessment distribution matrix, the abnormal index is determined and compared with the adaptive threshold calculated by the density clustering algorithm. If it is greater than the adaptive threshold, the fault diagnosis model is triggered. The fault-related data is collected in a targeted manner through the improved distributed tracing system. The fault-related data is constructed as a service call dependency graph based on the timing graph calculation algorithm combined with the gated graph attention mechanism. The abnormal propagation path is analyzed by a causal reasoning model based on a probability network in combination with the multi-dimensional feature vector, and a fault propagation chain is generated. The system log is obtained, and the system log and the fault propagation chain are associated with the deep semantic analysis model and the graph structure network. The context features are extracted and the fault feature map is generated in combination with the knowledge transfer technology. The fault feature map is multi-layered classified based on the integrated decision tree framework, the features of the abnormal propagation path are integrated, and the fault root cause analysis report and the risk level assessment result are output in combination with the hierarchical generalization strategy. The switching strategy parameter set is generated in combination with the dynamic programming algorithm, including:

[0025] The health score of each indicator in the system health assessment distribution matrix is ​​compared with the historical health data, the historical health data is grouped according to the indicator dimension, the distance between each data point and the adjacent data point is calculated to obtain the local data density, the local data density is sorted and the difference between the adjacent density values ​​is calculated, the density demarcation point is determined at the position where the difference jumps, the data is divided into multiple density intervals based on the density demarcation point, the highest density point is selected as the density center point in each density interval, the average distance from all data points in the density interval to the density center point is calculated as the characteristic radius of the current density interval, and the adaptive threshold is determined by the density weighting method of the density interval;

[0026] When the health score is lower than the adaptive threshold, a service call topology is constructed through a distributed tracing system, a tracing identifier is injected at the abnormal service node, the tracing identifier is transmitted along with the request on the call link, call data information is recorded for the request injected with the tracing identifier, full sampling is adopted for the node marked abnormal, and a sampling rate of 50% is adopted for the associated upstream and downstream nodes;

[0027] Aggregate call data with the same tracking identifier and construct a call sequence according to timestamps, perform semantic analysis on interface names in the call chain to calculate the similarity between interfaces, classify interfaces with high similarity into the same call group, analyze the call association features between the call groups, assign importance weights to each call group based on the call association features, and construct a weighted service call dependency graph based on the association degree between call groups and the importance weights;

[0028] The abnormal probability of the adjacent nodes when each service node in the service call dependency graph is abnormal is counted to establish a conditional probability relationship, and the abnormal state propagation sequence is analyzed in combination with the abnormal occurrence sequence of the indicators recorded in the multi-dimensional feature vector. The diffusion probability distribution of the abnormal event in the network is obtained through iterative calculation, and the abnormal propagation path is traced back to locate the abnormal source node. The main abnormal propagation path is screened based on the probability dependence strength between nodes to form a fault propagation chain;

[0029] Extract structured information fields from system logs, identify key information in log texts through word frequency analysis to build log templates, map log texts to the log templates and extract variable values ​​to capture changes in system operation status, perform vectorization on log texts to extract semantic features, convert topological structure information of the fault propagation chain into feature vectors, and fuse the semantic features with the feature vectors through an attention mechanism to form a fault context feature vector;

[0030] Retrieve cases with similar symptom characteristics to the current fault from the historical fault case library to extract fault handling experience, migrate historical fault modes to the current scenario through feature distribution adaptation, build a fault feature map based on the current fault manifestation, use a three-layer classification structure to locate the fault, perform coarse-grained classification based on symptom characteristics, refine the classification results based on the abnormal propagation path, and determine the fault type by integrating system status characteristics. Use the random forest method to integrate multiple classifier results at each layer of classification.

[0031] Based on the fault type, the current state of the system is determined and the system state vector is constructed. The impact of the switching operation on the system state is analyzed, the degree of system impact during the switching process is evaluated, and the switching path with the minimum system loss is obtained through iterative calculation of the state value, and a system switching strategy parameter set is generated.

[0032] In an optional embodiment,

[0033] Retrieve cases similar to the current fault symptom characteristics from the historical fault case library to extract fault handling experience, migrate historical fault modes to the current scenario through feature distribution adaptation, build a fault feature map based on the current fault manifestation, and locate the fault using a three-layer classification structure. Perform coarse-grained classification based on symptom characteristics, refine the classification results based on the abnormal propagation path, and determine the fault type by integrating system status characteristics. Each layer of classification uses the random forest method to integrate multiple classifiers. The results include:

[0034] Extract key features of historical fault cases, including system layer indicators, service layer indicators, component layer indicators and application layer indicators, segment the key features into time windows, the time window length is five minutes, the time window overlap time is one minute, calculate statistical feature values ​​within the time window, the statistical feature values ​​include mean, standard deviation, maximum value, minimum value, median, skewness, kurtosis and quantile, retain the original value for continuous features with a fixed value range, perform maximum and minimum value normalization on continuous features with inconsistent dimensions, and perform one-hot encoding conversion on discrete features;

[0035] Collect the monitoring data of the current fault in the last thirty minutes, obtain the current fault vector by performing feature extraction on the monitoring data, calculate the cosine similarity between the current fault vector and the historical fault case vector, select the historical fault cases with the cosine similarity greater than 0.8 as the candidate set, and reduce the cosine similarity threshold to 0.6 when the number of cases in the candidate set is less than ten;

[0036] Extracting environmental context information, the environmental context information includes hardware configuration information, deployment information and load feature information, constructing a feature graph, taking the key features as graph nodes, calculating the Pearson correlation coefficient between the graph nodes, establishing edges between the graph nodes when the absolute value of the Pearson correlation coefficient is greater than 0.6, taking the Pearson correlation coefficient as the edge weight, grouping the graph nodes using a hierarchical clustering method, setting a minimum distance threshold of 0.4, and using the graph edit distance to calculate the pattern similarity between the result of the hierarchical clustering and the combination of historical fault features in the candidate set;

[0037] A multi-layer random forest classifier is used for fault classification. The first-layer classifier divides faults into major categories based on system resource indicator subsets, service status indicator subsets, error log feature subsets, alarm information subsets and performance indicator subsets. The second-layer classifier subdivides faults into subcategories based on the timing relationship of abnormal nodes, propagation delay time, number of affected services and similarity of error types. The third-layer classifier determines the specific cause of the fault based on the system operation status curve, load change trend and resource usage fluctuation. The voting weight is determined according to the accuracy of each layer of classifiers on the validation set, and the predicted fault type is obtained through weighted voting.

[0038] In an optional embodiment,

[0039] Based on the risk level assessment result, it is determined whether to trigger the switch. If triggered, the intelligent switch execution module is started. The candidate nodes are screened from the global standby node pool through the consistent hashing algorithm combined with the virtual node mechanism and according to the switch policy parameter set. The standby node scoring system is constructed based on the characteristic indicators in the fault root cause analysis report based on the deep strategy learning algorithm. The decision process is optimized through the search tree and dynamic scoring is performed. The candidate node with the highest score value is used as the switching target node and the double buffer zero copy mechanism is initialized based on the fault propagation chain. The database operation request is prioritized in combination with the multi-level feedback queue algorithm. Sequential switching is performed based on the priority division result. The performance indicators in the switching process are compared and modeled with the system health assessment distribution matrix in combination with the timing prediction algorithm. The switching policy parameter set is optimized until the switching is completed, including:

[0040] Obtain the risk level assessment result, determine the number of affected service nodes in combination with the service call chain, set a weight value according to the importance of the service node in the business system, and the weight value ranges from zero to one. Count the service unavailability time caused by the fault, count the amount of new data after the most recent data backup, calculate the impact score of the service node, the service unavailability time score, and the data loss risk score according to the preset weight ratio to obtain a fault risk score, and determine whether to trigger the switching operation based on the fault risk score;

[0041] If the switching operation is triggered, an integer hash ring is established through a consistent hashing algorithm, a hash value of the physical node identifier is calculated using a message digest algorithm to determine the position of the physical node on the hash ring, an increasing serial number is appended to each physical node identifier to generate a virtual node identifier, a hash value of the virtual node identifier is calculated to determine the position of the virtual node on the hash ring, and the virtual node hash value is stored in a jump list to obtain the candidate node;

[0042] Based on the candidate nodes, a state feature vector including current load indicators, resource usage, and network status is constructed through a deep strategy learning algorithm, the selection combination of the candidate nodes is used as the action space, the reward value is calculated based on the degree of improvement in system performance after switching, a sample pool of preset capacity is maintained from historical switching records, and training samples are randomly extracted for batch update;

[0043] Initialize the decision tree and record the candidate nodes, performance indicator values, number of alternatives and cumulative path scores at the decision tree nodes, select the candidate with the highest expected score for node expansion based on the scoring system, perform random sampling simulation on each layer of the decision tree and calculate the average score as the node score, and determine the optimal switching path according to the preset depth and branch number limit;

[0044] Based on the optimal switching path, a double buffer zero copy mechanism is started, two buffers of the same size are respectively set at the source node and the switching target node, the source node writes into the first buffer and transmits the second buffer data to the switching target node through direct memory access, when the first buffer is full, the data of the first buffer is switched to the second buffer to be written and transmitted, and the switching target node uses the receiving buffer and the persistent buffer to process the transmitted data alternately;

[0045] Based on the double-buffered zero-copy mechanism, multiple priority queues are constructed using a ring buffer, and the time slice size of each priority queue is set according to a preset ratio. When the high-priority queue is empty, the database operation requests in the low-priority queue are processed, and the requests that have exceeded the time slice and have not been processed are moved to the next-level queue, and the lock-free queue concurrency control is completed through atomic operations;

[0046] The size and step size of the performance monitoring sliding window are set, and the mean, variance, coefficient of variation and trend slope of the performance indicators in the sliding window are calculated. When the response time or error rate in the performance indicators exceeds the preset threshold in the system health assessment distribution matrix, the batch size and concurrency in the switching strategy parameter set are adjusted according to the preset ratio until the performance indicators are maintained within the range defined by the system health assessment distribution matrix.

[0047] In an optional embodiment,

[0048] If the switching operation is triggered, an integer hash ring is established by a consistent hashing algorithm, a hash value of a physical node identifier is calculated by a message digest algorithm to determine the position of the physical node on the hash ring, an increasing sequence number is added to each physical node identifier to generate a virtual node identifier, a hash value of the virtual node identifier is calculated to determine the position of the virtual node on the hash ring, and the virtual node hash value is stored in a jump table to obtain the candidate node, including:

[0049] Initialize an integer hash ring, wherein the hash ring is stored in a skip table structure, wherein a node of the skip table includes a hash value field and a node information field, wherein the node information field stores a physical node identifier and physical node status information, wherein the physical node status information includes a node load value, an amount of available resources, and a network connection status;

[0050] Read a physical node identifier from a global standby node pool, the physical node identifier including a network address and a port number, calculate a first hash value for the physical node identifier using a message digest algorithm, intercept the first half of the first hash value as a physical node position value, and store the physical node position value and the corresponding physical node identifier in the skip list structure;

[0051] Adding an increasing sequence number after the physical node identifier to generate multiple virtual node identifiers, the increasing sequence number is connected to the physical node identifier through a separator, using the message digest algorithm to calculate the virtual node identifier to obtain a virtual node hash value, intercepting the first half of the virtual node hash value as a virtual node position value, and storing the virtual node position value and the corresponding physical node information in the skip list structure;

[0052] Constructing a multi-level index linked list, wherein the nodes of each level of the index linked list in the multi-level index linked list are arranged in ascending order according to the virtual node position value, and the nodes of the index linked lists at adjacent levels establish jump pointers according to a preset probability, and the jump pointers are used to quickly locate the target interval between the multi-level index linked lists;

[0053] Receive a data migration request, the data migration request including a target data key value, calculate a target hash value of the target data key value using the message digest algorithm, search for a first virtual node greater than or equal to the target hash value starting from the highest-level index linked list in the multi-level index linked list, and use a physical node corresponding to the first virtual node as a target switching node;

[0054] A data transmission channel is established according to the physical node identifier of the target switching node, data on the source node is migrated to the target switching node according to the interval range of the virtual node position value, and the node status information in the jump table structure is updated.

[0055] A second aspect of an embodiment of the present invention provides an HBase client active / standby switching system based on fault perception, including:

[0056] The first unit is used to collect system operation data through the distributed fault monitoring module in the client and form a multi-dimensional system indicator data set through monitoring probes deployed in different network areas, perform data distribution learning on the multi-dimensional system indicator data set through an adaptive clustering algorithm and establish an anomaly detection baseline model, perform a sliding time window operation on the multi-dimensional system indicator data set based on the anomaly detection baseline model and the exponential moving average algorithm in combination with a low-pass filter to obtain pre-processed standard indicator data, perform time-frequency domain feature extraction on the standard indicator data through a multi-scale analysis method based on wavelet packet decomposition, perform dimensionality reduction in combination with a principal component analysis algorithm to obtain a multi-dimensional feature vector, and generate a system health assessment distribution matrix based on a preset deep neural network model and the anomaly detection baseline model;

[0057] The second unit is used to determine the abnormal index based on the system health assessment distribution matrix and compare it with the adaptive threshold calculated by the density clustering algorithm. If it is greater than the adaptive threshold, the fault diagnosis model is triggered, and the fault-related data is collected in a targeted manner through the improved distributed tracing system. The fault-related data is constructed as a service call dependency graph based on the timing graph calculation algorithm combined with the gated graph attention mechanism. The abnormal propagation path is analyzed by a causal reasoning model based on a probability network in combination with the multi-dimensional feature vector, and a fault propagation chain is generated. The system log is obtained, and the system log and the fault propagation chain are associated with the deep semantic analysis model and the graph structure network. The context features are extracted and a fault feature map is generated in combination with the knowledge transfer technology. The fault feature map is multi-layered classified based on the integrated decision tree framework, the features of the abnormal propagation path are integrated, and the fault root cause analysis report and risk level assessment results are output in combination with the hierarchical generalization strategy. The switching strategy parameter set is generated in combination with the dynamic programming algorithm;

[0058] The third unit is used to determine whether to trigger switching based on the risk level assessment result, and if triggered, start the intelligent switching execution module, screen candidate nodes from the global standby node pool through the consistent hashing algorithm combined with the virtual node mechanism and according to the switching policy parameter set, construct a standby node scoring system based on the characteristic indicators in the fault root cause analysis report based on the deep strategy learning algorithm, optimize the decision process through the search tree and perform dynamic scoring, use the candidate node with the highest score value as the switching target node and initialize the double buffer zero copy mechanism based on the fault propagation chain, prioritize database operation requests in combination with the multi-level feedback queue algorithm, perform sequential switching based on the priority division result, compare and model the performance indicators in the switching process with the system health assessment distribution matrix in combination with the timing prediction algorithm, and optimize the switching policy parameter set until the switching is completed.

[0059] According to a third aspect of the embodiments of the present invention,

[0060] An electronic device is provided, comprising:

[0061] processor;

[0062] a memory for storing processor-executable instructions;

[0063] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0064] According to a fourth aspect of the embodiments of the present invention,

[0065] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.

[0066] In the present invention, multi-dimensional system indicator data is collected through a distributed fault monitoring module, and an anomaly detection model is established in combination with adaptive clustering and deep neural networks, which realizes accurate identification and timely warning of system anomalies, improves the accuracy and real-time performance of fault detection, and effectively reduces the false alarm rate. An improved distributed tracing system and a causal reasoning model based on a probabilistic network are adopted, combined with deep semantic analysis and knowledge transfer technology, to construct a complete fault feature map, realize accurate positioning of the fault propagation path and rapid identification of the root cause, and improve the efficiency and accuracy of fault diagnosis. A spare node scoring system is established based on a deep strategy learning algorithm, and intelligent switching is realized in combination with a double-buffered zero-copy mechanism and a multi-level feedback queue algorithm. The smoothness of the switching process is ensured by dynamically optimizing the switching strategy parameters, which significantly improves the system availability and service continuity, and reduces the impact of the switching process on the business. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Figure 1 A schematic diagram of a process flow of a method for switching between active and standby HBase clients based on fault perception according to an embodiment of the present invention;

[0068] Figure 2 The present invention is a schematic diagram of the structure of an HBase client active-standby switching system based on fault perception according to an embodiment of the present invention. DETAILED DESCRIPTION

[0069] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0070] The technical solution of the present invention is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.

[0071] Figure 1 FIG. 1 is a flow chart of a method for switching between active and standby HBase clients based on fault perception according to an embodiment of the present invention. Figure 1 As shown, the method includes:

[0072] S1. Collect system operation data through the distributed fault monitoring module in the client and form a multi-dimensional system indicator data set through monitoring probes deployed in different network areas, perform data distribution learning on the multi-dimensional system indicator data set through an adaptive clustering algorithm and establish an anomaly detection baseline model, perform a sliding time window operation on the multi-dimensional system indicator data set based on the anomaly detection baseline model and the exponential moving average algorithm in combination with a low-pass filter to obtain pre-processed standard indicator data, extract time-frequency domain features of the standard indicator data through a multi-scale analysis method based on wavelet packet decomposition, and reduce the dimension in combination with the principal component analysis algorithm to obtain a multi-dimensional feature vector, and generate a system health assessment distribution matrix based on a pre-set deep neural network model and the anomaly detection baseline model;

[0073] The adaptive clustering algorithm is an algorithm that dynamically adjusts the clustering process. It adaptively updates the cluster center and the number of categories according to changes in data distribution or clustering targets. It is suitable for processing dynamic and non-stationary data sets and is often used in time series analysis and online learning tasks. The anomaly detection baseline model is a reference model for identifying abnormal patterns in data sets. Based on the baseline characteristics of normal behavior, it uses statistical methods, machine learning or deep learning techniques to determine whether the input data deviates from the baseline range, thereby detecting anomalies. The exponential moving average algorithm is a weighted method for calculating the average value, which gives higher weights to the most recent observations. The multiscale analysis method based on wavelet packet decomposition is a signal processing technology that analyzes the multiscale characteristics of the signal by decomposing the signal into multiple frequency subbands. The principal component analysis algorithm is a dimensionality reduction technology that aims to project high-dimensional data into a low-dimensional space through linear transformation. The system health assessment distribution matrix is ​​a matrix form model used to represent and evaluate the health status of the system. The multidimensional status indicators of the system are mapped into a matrix, and the health distribution is analyzed in combination with statistical or machine learning methods to provide support for system performance prediction and fault diagnosis.

[0074] In an optional embodiment,

[0075] The distributed fault monitoring module in the client collects system operation data and forms a multi-dimensional system indicator data set through monitoring probes deployed in different network areas. The adaptive clustering algorithm is used to learn the data distribution of the multi-dimensional system indicator data set and establish an anomaly detection baseline model, including:

[0076] The distributed fault monitoring module collects system operation data, and the system operation data includes processor utilization index, memory occupancy index, disk performance index, and network performance index. The processor utilization index is obtained by collecting the user state utilization rate, system state utilization rate, input and output waiting rate, and idle rate of each processor core. The memory occupancy index is obtained by collecting the physical memory usage, virtual memory usage, cache usage, and buffer usage. The disk performance index is obtained by collecting the read and write rate, average response time, input and output queue length, and device utilization of each disk device. The network performance index is obtained by collecting the input and output bandwidth utilization, network delay, packet loss rate, and retransmission rate.

[0077] The distributed fault monitoring module collects the system operation data through monitoring probes deployed in different network areas to form a multi-dimensional system indicator data set, wherein the monitoring probes are composed of a first-level probe node, a second-level probe node, and a third-level probe node. The first-level probe node is configured according to the ratio of the number of racks and is responsible for local data collection. The second-level probe node is configured according to the number of computer rooms and is responsible for aggregating the data collected by the first-level probe node for preprocessing. The third-level probe node is configured after dividing the area according to the geographical location and network topology structure and is responsible for coordinating and managing the data collection tasks in the area. The first-level probe node, the second-level probe node, and the third-level probe node maintain connection through a heartbeat mechanism.

[0078] After the multidimensional system indicator data set is standardized and converted, the Euclidean distance between the sample and the cluster center is calculated based on the initial cluster center for classification. When the variance of the samples in the cluster exceeds the preset threshold, a new cluster center is added in the direction of the maximum variance in the current cluster. When the distance between adjacent cluster centers is less than the preset threshold, adjacent clusters are merged. If the moving distance of the cluster center is less than the preset convergence threshold for 100 consecutive iterations, the data distribution learning is considered to be completed.

[0079] Based on the results of data distribution learning, the mean and standard deviation of samples in each cluster are calculated, and the multiple range of the mean and the standard deviation is set as the normal value range of the current indicator. At the same time, the historical data is divided into hourly cycles to establish an hourly baseline, a daily baseline is established according to a daily cycle, and a monthly baseline is established according to a monthly cycle. The hourly baseline, the daily baseline, and the monthly baseline are weighted to form the anomaly detection baseline model.

[0080] The Euclidean distance is a commonly used distance measurement method, which is used to calculate the straight-line distance between two points in space. The initial cluster center is a set of center points preset at the beginning of the clustering algorithm, which is used to determine the initial position of the data grouping. The selection of the initial cluster center has an important influence on the convergence speed of the algorithm and the quality of the results. It is usually selected by random initialization, k-means++ or heuristic methods. The sample variance is a statistic that measures the degree of dispersion of a set of sample data distribution.

[0081] The implementation of the distributed fault monitoring system first deploys a monitoring probe network. The monitoring probe adopts a three-level architecture design: the first-level probe nodes are deployed at the rack level, and are deployed according to the principle of configuring at least one probe node for each rack; the second-level probe nodes are deployed at the computer room level, and 2-5 nodes are configured according to the scale of the computer room; the third-level probe nodes are deployed according to geographical areas, usually 3 nodes are configured in each area to form redundant backup, and the probe nodes maintain heartbeats through TCP long connections with a heartbeat interval of 5 seconds. If the heartbeat timeouts for three consecutive times, the node is considered offline.

[0082] During the system operation data collection process, the first-level probe node collects basic indicator data every 10 seconds. The processor utilization indicator collection includes four dimensions: user-state CPU utilization, system-state CPU utilization, IO waiting rate and idle rate; memory indicator collection includes physical memory usage, virtual memory usage, page cache usage and buffer usage; disk performance indicator collection includes read and write rate (unit MB / s), average response time (unit ms), IO queue length and device utilization; network performance indicator collection includes inbound and outbound bandwidth utilization, network delay (unit ms), packet loss rate (percentage) and TCP retransmission rate.

[0083] The secondary probe node obtains data from the primary node every minute and performs preprocessing, including outlier filtering and data standardization. The 3-sigma principle is used for outlier filtering, and data that deviates from the mean by more than 3 standard deviations is marked as abnormal. Data standardization uses the minimum-maximum normalization method to convert each indicator to the 0-1 interval. The tertiary probe node obtains processed data from the secondary node every 5 minutes, and dynamically adjusts the collection frequency of nodes at all levels according to the actual load conditions.

[0084] The adaptive clustering algorithm first selects the initial cluster center for the standardized multidimensional indicator data. The selection method is to randomly select 10 sample points in the data space as the initial cluster center. For each data sample, the Euclidean distance between it and all cluster centers is calculated, and the sample is divided into the cluster with the closest distance. When the sample variance within a cluster exceeds 0.3, a new cluster center is added on the feature dimension with the largest variance; when the distance between adjacent cluster centers is less than 0.1, the two clusters are merged. During the iteration process, if the moving distance of the cluster center is less than 0.001 for 100 consecutive iterations, the clustering is considered to have converged.

[0085] When building a baseline model, first calculate the mean vector and standard deviation vector of each cluster. Taking CPU usage as an example, if the mean of a cluster is 60% and the standard deviation is 10%, set 50%-70% as the normal value range. When building a time series baseline, divide the historical data into 1-hour, 24-hour, and 30-day periods, and build corresponding baseline models respectively. The final anomaly detection baseline combines the baselines of the three time scales in a weighted manner, with weights of 0.5 for the hourly baseline, 0.3 for the daily baseline, and 0.2 for the monthly baseline.

[0086] In this embodiment, efficient data collection in a large-scale distributed environment is achieved through a three-level probe architecture, which ensures the real-time and reliability of data while reducing network overhead and storage costs. An adaptive clustering algorithm is used to realize automatic learning of the distribution of system indicator data. The algorithm can dynamically adjust the number of clusters according to data characteristics, thereby improving the accuracy and adaptability of anomaly detection. The combination of multi-time scale baseline models takes into account the changing laws of system behavior in different time periods, and can effectively identify short-term fluctuations and long-term trend changes, thereby improving the robustness of anomaly detection.

[0087] In an optional embodiment,

[0088] Based on the anomaly detection baseline model and the exponential moving average algorithm, a sliding time window operation is performed on the multidimensional system indicator data set in combination with a low-pass filter to obtain preprocessed standard indicator data, time-frequency domain feature extraction is performed on the standard indicator data through a multi-scale analysis method based on wavelet packet decomposition, and a multi-dimensional feature vector is obtained by dimensionality reduction in combination with a principal component analysis algorithm, and a system health assessment distribution matrix is ​​generated based on a pre-set deep neural network model and the anomaly detection baseline model, including:

[0089] The distributed fault monitoring module performs a sliding time window operation on the multidimensional system indicator data set based on the anomaly detection baseline model and the exponential moving average algorithm in combination with a low-pass filter, fills the continuous indicator with linear interpolation of the two normal values ​​before and after, and fills the discrete indicator with the nearest neighbor normal value, distinguishes high-frequency change indicators from low-frequency change indicators based on the ratio of the standard deviation to the mean, and smoothes the high-frequency change indicators and the low-frequency change indicators with different smoothing factors to obtain pre-processed standard indicator data;

[0090] The time-frequency domain features of the standard indicator data are extracted by a multi-scale analysis method based on wavelet packet decomposition, the standard indicator data is decomposed into multiple frequency bands, the energy value of each frequency band is calculated, the frequency bands whose energy proportion is lower than a preset energy proportion threshold are eliminated, the mean, variance, skewness, kurtosis, maximum value and minimum value of the retained frequency bands are extracted as statistical features, and the energy ratio between different frequency bands is calculated based on the statistical features to obtain the time-frequency domain features;

[0091] The distributed fault monitoring module calculates the covariance matrix of the feature matrix based on the principal component analysis algorithm to obtain eigenvalues ​​and eigenvectors, selects eigenvectors corresponding to the eigenvalues ​​based on a preset variance contribution rate threshold to construct a conversion matrix, and maps the time-frequency domain features to the principal component space to obtain a multi-dimensional eigenvector;

[0092] The distributed fault monitoring module inputs the multi-dimensional feature vector into a preset deep neural network model, wherein the number of input layer nodes of the deep neural network model is the same as the number of principal components, the number of hidden layer nodes is set according to the ratio of the input dimensions, the number of output layer nodes is the same as the type of monitoring indicators, the parameters of the deep neural network model are trained based on normal state data, the multi-dimensional feature vector is input into the trained deep neural network model to obtain a health score, and normalized in combination with the anomaly detection baseline model to generate a system health assessment distribution matrix.

[0093] The nearest neighbor normal value is an approximate value of a sample within a normal range estimated based on the nearest neighbor method, and is usually used for anomaly detection and data repair. By calculating the distance between the sample and the nearest neighbor point, the degree of deviation of the sample can be determined and a reference value can be obtained. The variance contribution rate threshold is an indicator used to determine the number of principal components selected in principal component analysis. It is defined as the cumulative variance contribution rate of the selected principal component exceeding the set threshold, which is used to ensure that sufficient information is retained while reducing the data dimension.

[0094] Preprocess the multidimensional system indicator data set. Process the raw data through sliding time window operation. The window size can be set to 30 minutes and the step size is 1 minute. For missing values ​​in continuous indicators, linear interpolation method is used to fill in. For example, if the temperature value at a certain moment is missing, the two normal temperature values ​​before and after are taken for linear interpolation to obtain the filling value. For discrete indicators such as device status, the nearest neighbor normal value is used for filling. For example, when the device status value is missing, the most recent normal status value is used for filling.

[0095] The ratio of the standard deviation to the mean of each indicator was calculated, and those with a ratio greater than 0.5 were classified as high-frequency change indicators, and those less than 0.5 were classified as low-frequency change indicators. The high-frequency change indicators were processed with an exponential moving average using a smoothing factor of 0.3, and the low-frequency change indicators were processed with a smoothing factor of 0.7. At the same time, a low-pass filter was introduced to eliminate high-frequency noise, and the cutoff frequency was set to one-tenth of the original signal frequency.

[0096] Perform wavelet packet decomposition on the preprocessed standard indicator data and decompose the signal into 8 frequency bands. Calculate the energy value of each frequency band, and remove the frequency bands with energy ratio less than 5%. Extract statistical features from the retained frequency bands, including mean, variance, skewness, kurtosis, maximum value and minimum value. Taking a temperature sensor data as an example, the low frequency band energy ratio after decomposition is 60%, the mid-frequency band is 30%, and the high frequency band is 10%, so the low frequency and mid-frequency bands are retained. Calculate the energy ratio of these two frequency bands as the time-frequency domain feature.

[0097] Based on principal component analysis, feature dimension reduction is performed and the covariance matrix of the feature matrix is ​​calculated. The variance contribution rate threshold is set to 85%, and the eigenvectors whose cumulative variance contribution rate reaches the threshold are selected to construct the transformation matrix. The time-frequency domain features are mapped to the principal component space to obtain the eigenvector after dimension reduction.

[0098] Construct a deep neural network model. The number of nodes in the input layer is the same as the number of principal components. If the number of principal components is 10, the input layer has 10 nodes. The hidden layer adopts a three-layer structure, and the number of nodes is 2 times, 1.5 times, and 1 times the input dimension respectively. The number of nodes in the output layer is the same as the type of monitoring indicators. Use normal operating status data to train the model, and use the back propagation algorithm to optimize the network parameters. Input the feature vector into the trained model to obtain the health score, and combine the anomaly detection baseline model to normalize the score to the range of 0-1 to generate the system health assessment distribution matrix.

[0099] In this embodiment, through sliding time window and data filling preprocessing, combined with exponential moving average and low-pass filtering, the noise and abnormal fluctuations in the original data are effectively eliminated, and the accuracy and reliability of subsequent analysis are improved. Wavelet packet decomposition is used for multi-scale analysis, time-frequency domain features are extracted, and principal component analysis is combined for dimensionality reduction, which not only retains the key feature information of the signal, but also reduces data redundancy and improves the efficiency of feature extraction. System health status assessment is implemented based on deep neural networks, and the correlation between features is fully explored through multi-layer nonlinear mapping. The evaluation results are standardized in combination with the anomaly detection baseline model, which improves the accuracy and interpretability of fault diagnosis.

[0100] S2. Determine the abnormal index based on the system health assessment distribution matrix and compare it with the adaptive threshold calculated by the density clustering algorithm. If it is greater than the adaptive threshold, trigger the fault diagnosis model, collect fault-related data through the improved distributed tracing system, construct the fault-related data into a service call dependency graph based on the timing graph calculation algorithm combined with the gated graph attention mechanism, analyze the abnormal propagation path through the causal reasoning model based on the probability network in combination with the multi-dimensional feature vector, generate the fault propagation chain, obtain the system log, perform correlation analysis on the system log and the fault propagation chain based on the deep semantic analysis model and the graph structure network, extract context features and generate the fault feature map in combination with the knowledge transfer technology, perform multi-layer classification operation on the fault feature map based on the integrated decision tree framework, integrate the features of the abnormal propagation path, output the fault root cause analysis report and risk level assessment results in combination with the hierarchical generalization strategy, and generate the switching strategy parameter set in combination with the dynamic programming algorithm;

[0101] The density clustering algorithm is an unsupervised learning algorithm that clusters based on the density distribution of sample points. It can identify clusters of any shape and effectively process noise points. The time series graph calculation algorithm is an algorithm for processing graph data with time information, analyzing the dynamic change trend of the relationship between nodes, and is often used in dynamic network modeling and prediction tasks. The causal reasoning model based on probability network uses probabilistic graph models (such as Bayesian networks or Markov networks) for causal reasoning, and realizes the quantification and analysis of causal relationships through conditional independence assumptions and probability updates. The deep semantic analysis model is a natural language understanding model based on deep learning technology, which is used to capture deep semantic information in text. Information is usually combined with an embedding layer and an attention mechanism and is applied to tasks such as semantic classification and sentiment analysis. The knowledge transfer technology is a method that uses pre-trained models or existing knowledge to improve the performance of new tasks. Through transfer learning, the generalization ability of the model can be improved when the amount of data is insufficient. The integrated decision tree framework is a machine learning method that improves prediction accuracy by combining multiple decision tree models (such as random forests or gradient boosting trees), which can effectively reduce the overfitting risk of a single model. The hierarchical generalization strategy is a method that gradually generalizes the features and decision-making capabilities of the model through a hierarchical structure, abstracts and optimizes the input data at different levels, and is suitable for staged processing of complex tasks.

[0102] In an optional embodiment,

[0103] Based on the system health assessment distribution matrix, the abnormal index is determined and compared with the adaptive threshold calculated by the density clustering algorithm. If it is greater than the adaptive threshold, the fault diagnosis model is triggered. The fault-related data is collected in a targeted manner through the improved distributed tracing system. The fault-related data is constructed as a service call dependency graph based on the timing graph calculation algorithm combined with the gated graph attention mechanism. The abnormal propagation path is analyzed by a causal reasoning model based on a probability network in combination with the multi-dimensional feature vector, and a fault propagation chain is generated. The system log is obtained, and the system log and the fault propagation chain are associated with the deep semantic analysis model and the graph structure network. The context features are extracted and the fault feature map is generated in combination with the knowledge transfer technology. The fault feature map is multi-layered classified based on the integrated decision tree framework, the features of the abnormal propagation path are integrated, and the fault root cause analysis report and the risk level assessment result are output in combination with the hierarchical generalization strategy. The switching strategy parameter set is generated in combination with the dynamic programming algorithm, including:

[0104] The health score of each indicator in the system health assessment distribution matrix is ​​compared with the historical health data, the historical health data is grouped according to the indicator dimension, the distance between each data point and the adjacent data point is calculated to obtain the local data density, the local data density is sorted and the difference between the adjacent density values ​​is calculated, the density demarcation point is determined at the position where the difference jumps, the data is divided into multiple density intervals based on the density demarcation point, the highest density point is selected as the density center point in each density interval, the average distance from all data points in the density interval to the density center point is calculated as the characteristic radius of the current density interval, and the adaptive threshold is determined by the density weighting method of the density interval;

[0105] When the health score is lower than the adaptive threshold, a service call topology is constructed through a distributed tracing system, a tracing identifier is injected at the abnormal service node, the tracing identifier is transmitted along with the request on the call link, call data information is recorded for the request injected with the tracing identifier, full sampling is adopted for the node marked abnormal, and a sampling rate of 50% is adopted for the associated upstream and downstream nodes;

[0106] Aggregate call data with the same tracking identifier and construct a call sequence according to timestamps, perform semantic analysis on interface names in the call chain to calculate the similarity between interfaces, classify interfaces with high similarity into the same call group, analyze the call association features between the call groups, assign importance weights to each call group based on the call association features, and construct a weighted service call dependency graph based on the association degree between call groups and the importance weights;

[0107] The abnormal probability of the adjacent nodes when each service node in the service call dependency graph is abnormal is counted to establish a conditional probability relationship, and the abnormal state propagation sequence is analyzed in combination with the abnormal occurrence sequence of the indicators recorded in the multi-dimensional feature vector. The diffusion probability distribution of the abnormal event in the network is obtained through iterative calculation, and the abnormal propagation path is traced back to locate the abnormal source node. The main abnormal propagation path is screened based on the probability dependence strength between nodes to form a fault propagation chain;

[0108] Extract structured information fields from system logs, identify key information in log texts through word frequency analysis to build log templates, map log texts to the log templates and extract variable values ​​to capture changes in system operation status, perform vectorization on log texts to extract semantic features, convert topological structure information of the fault propagation chain into feature vectors, and fuse the semantic features with the feature vectors through an attention mechanism to form a fault context feature vector;

[0109] Retrieve cases with similar symptom characteristics to the current fault from the historical fault case library to extract fault handling experience, migrate historical fault modes to the current scenario through feature distribution adaptation, build a fault feature map based on the current fault manifestation, use a three-layer classification structure to locate the fault, perform coarse-grained classification based on symptom characteristics, refine the classification results based on the abnormal propagation path, and determine the fault type by integrating system status characteristics. Use the random forest method to integrate multiple classifier results at each layer of classification.

[0110] Based on the fault type, the current state of the system is determined and the system state vector is constructed. The impact of the switching operation on the system state is analyzed, the degree of system impact during the switching process is evaluated, and the switching path with the minimum system loss is obtained through iterative calculation of the state value, and a system switching strategy parameter set is generated.

[0111] The density demarcation point is an important parameter used to distinguish core points from boundary points in the density clustering algorithm. Its value is calculated based on the local point density and determines the boundary shape of the cluster. The tracking identifier is a unique tag assigned to each data object or transaction, which is used to achieve efficient tracking and tagging in dynamic systems. The full sampling is a data sampling method, which refers to extracting all samples from the population for analysis and modeling. It is usually suitable for scenarios with moderate sample size and high precision requirements. The service call dependency graph is a directed graph that reflects the call relationship between services. It shows the interaction path and dependency relationship between services in a microservice system and is often used to optimize the service architecture.

[0112] During the construction of the system health assessment distribution matrix, each monitoring indicator is collected and evaluated in real time. Taking the service response time as an example, assuming that the response time of a microservice under normal conditions is 100ms, when the response time is observed to exceed 200ms, an abnormality judgment needs to be made. Through density clustering, historical data is divided into multiple density intervals according to the response time distribution characteristics, such as 50-100ms, 100-150ms, 150-200ms, etc. In each interval, the point with the highest data density is selected as the center, and the average distance from all data points in the interval to the center point is calculated as the characteristic radius. By calculating the weights of different density intervals, the adaptive threshold of the response time is finally obtained, such as 180ms.

[0113] When an anomaly is detected, the improved distributed tracing system will inject a tracing identifier at the abnormal service. Taking the order processing service as an example, if an abnormal order processing delay is found, a unique tracing ID, such as "trace-order-123", will be injected into the service node. This tracing identifier will be passed to the associated payment service, inventory service and other nodes along with the request call. Full data collection is performed on the request path where the tracing identifier is injected, including call time, call parameters, return results and other information. A 50% sampling rate is used for data collection for the associated upstream and downstream service nodes.

[0114] Based on the collected call data, the service call dependency is constructed through the timing graph calculation algorithm. For example, analysis finds that the order processing service will call the payment verification interface and the inventory query interface. These interfaces may be classified into the same call group based on semantic similarity. By analyzing the call frequency, call timing and other features, weights are assigned to different call groups, and finally a weighted service call dependency graph is formed.

[0115] In the abnormal propagation analysis phase, the probability of abnormal state propagation between service nodes is counted. For example, when an abnormality occurs in the order processing service, the probability of abnormality in the associated payment service is 0.8, and the probability of abnormality in the inventory service is 0.3. Combined with the occurrence order of system indicator abnormalities, the diffusion path of the abnormality in the entire service network is obtained through iterative calculation, and finally the abnormal source node is located and the fault propagation chain is generated.

[0116] During the system log analysis process, extract key information from the log. For example, from the "Order processing timeout: order_id=12345, processing_time=250ms" log, extract structured information such as operation type, order ID, and processing time. Build a log template through word frequency statistics, and replace specific parameter values ​​with variable placeholders. Combine the topological structure characteristics of the fault propagation chain to form a complete fault context feature.

[0117] In the fault diagnosis phase, similar cases are retrieved from the historical fault database. For example, if the current fault manifests as an order processing timeout, the system will retrieve similar timeout fault cases in history, extract processing experience, and build a fault feature map based on the current scenario characteristics. A three-layer classification structure is used to locate faults, and rough classification is performed based on timeouts, errors, and other phenomena. The abnormal propagation path is combined for subdivision, and the specific fault type is determined by integrating the system status. A switching strategy is generated based on the fault diagnosis results. The impact of different switching schemes on the system is evaluated, the switching path with the least system loss is selected, and the specific switching parameter configuration is output.

[0118] In this embodiment, adaptive thresholds and improved distributed tracing are used to improve the accuracy of anomaly detection, reduce false alarm rates, and achieve accurate fault data collection. Combined with multi-dimensional feature analysis and deep semantic understanding, fault propagation paths and root causes are accurately identified, improving the accuracy and efficiency of fault diagnosis. A switching strategy generation solution based on dynamic programming minimizes the impact of the switching process on the system while ensuring service availability, thereby improving the overall reliability of the system.

[0119] In an optional embodiment,

[0120] Retrieve cases similar to the current fault symptom characteristics from the historical fault case library to extract fault handling experience, migrate historical fault modes to the current scenario through feature distribution adaptation, build a fault feature map based on the current fault manifestation, and locate the fault using a three-layer classification structure. Perform coarse-grained classification based on symptom characteristics, refine the classification results based on the abnormal propagation path, and determine the fault type by integrating system status characteristics. Each layer of classification uses the random forest method to integrate multiple classifiers. The results include:

[0121] Extract key features of historical fault cases, including system layer indicators, service layer indicators, component layer indicators and application layer indicators, segment the key features into time windows, the time window length is five minutes, the time window overlap time is one minute, calculate statistical feature values ​​within the time window, the statistical feature values ​​include mean, standard deviation, maximum value, minimum value, median, skewness, kurtosis and quantile, retain the original value for continuous features with a fixed value range, perform maximum and minimum value normalization on continuous features with inconsistent dimensions, and perform one-hot encoding conversion on discrete features;

[0122] Collect the monitoring data of the current fault in the last thirty minutes, obtain the current fault vector by performing feature extraction on the monitoring data, calculate the cosine similarity between the current fault vector and the historical fault case vector, select the historical fault cases with the cosine similarity greater than 0.8 as the candidate set, and reduce the cosine similarity threshold to 0.6 when the number of cases in the candidate set is less than ten;

[0123] Extracting environmental context information, the environmental context information includes hardware configuration information, deployment information and load feature information, constructing a feature graph, taking the key features as graph nodes, calculating the Pearson correlation coefficient between the graph nodes, establishing edges between the graph nodes when the absolute value of the Pearson correlation coefficient is greater than 0.6, taking the Pearson correlation coefficient as the edge weight, grouping the graph nodes using a hierarchical clustering method, setting a minimum distance threshold of 0.4, and using the graph edit distance to calculate the pattern similarity between the result of the hierarchical clustering and the combination of historical fault features in the candidate set;

[0124] A multi-layer random forest classifier is used for fault classification. The first-layer classifier divides faults into major categories based on system resource indicator subsets, service status indicator subsets, error log feature subsets, alarm information subsets and performance indicator subsets. The second-layer classifier subdivides faults into subcategories based on the timing relationship of abnormal nodes, propagation delay time, number of affected services and similarity of error types. The third-layer classifier determines the specific cause of the fault based on the system operation status curve, load change trend and resource usage fluctuation. The voting weight is determined according to the accuracy of each layer of classifiers on the validation set, and the predicted fault type is obtained through weighted voting.

[0125] The Pearson correlation coefficient is a statistical indicator used to measure the linear correlation between two variables. Its value is between -1 and 1, indicating complete negative correlation and complete positive correlation, respectively, and 0 indicates no correlation. The multi-layer random forest classifier is a classification framework that organizes multiple random forest models in a hierarchical structure. The output of each layer can be used as the input of the next layer to further improve the classification performance and robustness.

[0126] In the process of distributed system fault diagnosis, it is first necessary to extract features and preprocess historical fault cases. System layer indicators include CPU usage, memory occupancy, disk IO read and write rates, and network packet sending and receiving rates; service layer indicators include request response time, throughput, number of concurrent connections, and queue backlog; component layer indicators include cache hit rate, connection pool usage, thread pool activity, and message queue accumulation; application layer indicators include business success rate, error rate, number of timeouts, and number of retries.

[0127] The collected monitoring data is segmented according to five-minute time windows, and the adjacent windows are overlapped for one minute to capture abnormal features across windows. The statistical characteristic values ​​of the indicators are calculated in each time window, including the mean reflecting the overall level of the indicator, the standard deviation reflecting the degree of fluctuation of the indicator, the maximum and minimum values ​​reflecting the range of the indicator, the median reflecting the central tendency of the indicator, the skewness reflecting the symmetry of the indicator distribution, the kurtosis reflecting the sharpness of the indicator distribution, and the quantile reflecting the distribution characteristics of the indicator.

[0128] For indicators such as CPU usage that range from 0 to 100, the original values ​​are retained. For indicators with inconsistent dimensions such as response time, normalization is performed to map them to the range of 0 to 1. For discrete features such as error types, they are converted into one-hot encoding. Through feature engineering, a standardized set of historical fault feature vectors is obtained.

[0129] When a new fault occurs in the system, the monitoring data of the 30 minutes before the fault occurs is collected, and the data is subjected to the same feature extraction and preprocessing as the historical cases to obtain the current fault vector. The similarity between the current fault vector and the historical case vector is calculated, and the cases with similarity higher than the threshold are selected as candidate sets. If the number of candidate set cases is less than ten, the similarity threshold is appropriately lowered to obtain more reference cases.

[0130] Extract environmental context information to build a feature association graph. Hardware configuration information includes the number of CPU cores, memory size, disk capacity, and network card bandwidth. Deployment information includes availability zones, cluster size, version number, and deployment architecture. Load characteristics include request mode, data level, access frequency, and business type. Use these features as graph nodes, calculate the correlation between features as edge weights, and cluster and group features using a community discovery algorithm.

[0131] Multi-layer classifiers are used for fault diagnosis. The first layer classifies faults into resource exhaustion, performance bottlenecks, component failures, etc. based on system resource usage; the second layer classifies faults into single point faults, cascading faults, avalanche effects, etc. based on abnormal propagation characteristics; the third layer determines the specific cause of the fault, such as memory leaks, deadlocks, network partitions, etc., based on system behavior patterns. The prediction results of each layer of classifiers are weighted voted to obtain the final diagnosis result.

[0132] Taking a cache system failure as an example, the system found that the response delay increased suddenly, the memory usage rate continued to rise, and the cache hit rate decreased. The first-level classifier determined it to be a performance failure based on resource indicators, the second-level classifier determined it to be a single point failure based on the abnormal propagation characteristics, and the third-level classifier determined that the root cause was a memory leak based on the memory growth curve and access pattern, which was consistent with the diagnosis results of similar failures in the historical case library.

[0133] In this embodiment, by extracting experience from historical fault cases and combining them with current fault features, possible fault causes can be quickly located, reducing the time and errors of manual analysis. Through feature distribution adaptation and fusion of environmental context information, fault conditions in different system environments can be better handled, and the reliability of diagnostic results is improved. The application of multi-layer random forest classifiers enables the system to automatically learn complex fault modes and make accurate judgments in new fault scenarios, greatly reducing the need for manual intervention.

[0134] S3. Determine whether to trigger switching based on the risk level assessment result. If triggered, start the intelligent switching execution module, screen candidate nodes from the global standby node pool through the consistent hashing algorithm combined with the virtual node mechanism and according to the switching policy parameter set, construct a standby node scoring system based on the characteristic indicators in the fault root cause analysis report based on the deep strategy learning algorithm, optimize the decision process through the search tree and perform dynamic scoring, use the candidate node with the highest score as the switching target node and initialize the double buffer zero copy mechanism based on the fault propagation chain, prioritize database operation requests in combination with the multi-level feedback queue algorithm, perform sequential switching based on the priority division results, compare and model the performance indicators in the switching process with the system health assessment distribution matrix in combination with the timing prediction algorithm, and optimize the switching policy parameter set until the switching is completed.

[0135] The consistent hashing algorithm is an efficient distributed hashing algorithm used to achieve load balancing and dynamic node expansion, and is often used in distributed cache and distributed storage systems. The virtual node mechanism is an optimization technology for the consistent hashing algorithm. By allocating multiple virtual nodes to physical nodes, the load distribution is smoothed, the amount of data migration is reduced, and the fault tolerance of the system is enhanced. The deep strategy learning algorithm is a method that combines deep learning and reinforcement learning. It is used to learn complex strategies to optimize long-term benefits in high-dimensional state space and is widely used in games, robots and dynamic programming. The search tree optimization decision is a method for achieving optimal decision-making by constructing and traversing search trees. It is often used in path planning, combinatorial optimization and game theory problems. The double buffer zero copy mechanism is an efficient memory data transmission technology that uses double buffers and memory mapping to achieve fast data exchange, avoid multiple copies of data between user state and kernel state, and improve transmission performance. The multi-level feedback queue algorithm is a scheduling algorithm that dynamically adjusts task priorities. It is widely used in process scheduling of operating systems to improve system throughput and fairness.

[0136] In an optional embodiment,

[0137] Based on the risk level assessment result, it is determined whether to trigger the switch. If triggered, the intelligent switch execution module is started. The candidate nodes are screened from the global standby node pool through the consistent hashing algorithm combined with the virtual node mechanism and according to the switch policy parameter set. The standby node scoring system is constructed based on the characteristic indicators in the fault root cause analysis report based on the deep strategy learning algorithm. The decision process is optimized through the search tree and dynamic scoring is performed. The candidate node with the highest score value is used as the switching target node and the double buffer zero copy mechanism is initialized based on the fault propagation chain. The database operation request is prioritized in combination with the multi-level feedback queue algorithm. Sequential switching is performed based on the priority division result. The performance indicators in the switching process are compared and modeled with the system health assessment distribution matrix in combination with the timing prediction algorithm. The switching policy parameter set is optimized until the switching is completed, including:

[0138] Obtain the risk level assessment result, determine the number of affected service nodes in combination with the service call chain, set a weight value according to the importance of the service node in the business system, and the weight value ranges from zero to one. Count the service unavailability time caused by the fault, count the amount of new data after the most recent data backup, calculate the impact score of the service node, the service unavailability time score, and the data loss risk score according to the preset weight ratio to obtain a fault risk score, and determine whether to trigger the switching operation based on the fault risk score;

[0139] If the switching operation is triggered, an integer hash ring is established through a consistent hashing algorithm, a hash value of the physical node identifier is calculated using a message digest algorithm to determine the position of the physical node on the hash ring, an increasing serial number is appended to each physical node identifier to generate a virtual node identifier, a hash value of the virtual node identifier is calculated to determine the position of the virtual node on the hash ring, and the virtual node hash value is stored in a jump list to obtain the candidate node;

[0140] Based on the candidate nodes, a state feature vector including current load indicators, resource usage, and network status is constructed through a deep strategy learning algorithm, the selection combination of the candidate nodes is used as the action space, the reward value is calculated based on the degree of improvement in system performance after switching, a sample pool of preset capacity is maintained from historical switching records, and training samples are randomly extracted for batch update;

[0141] Initialize the decision tree and record the candidate nodes, performance indicator values, number of alternatives and cumulative path scores at the decision tree nodes, select the candidate with the highest expected score for node expansion based on the scoring system, perform random sampling simulation on each layer of the decision tree and calculate the average score as the node score, and determine the optimal switching path according to the preset depth and branch number limit;

[0142] Based on the optimal switching path, a double buffer zero copy mechanism is started, two buffers of the same size are respectively set at the source node and the switching target node, the source node writes into the first buffer and transmits the second buffer data to the switching target node through direct memory access, when the first buffer is full, the data of the first buffer is switched to the second buffer to be written and transmitted, and the switching target node uses the receiving buffer and the persistent buffer to process the transmitted data alternately;

[0143] Based on the double-buffered zero-copy mechanism, multiple priority queues are constructed using a ring buffer, and the time slice size of each priority queue is set according to a preset ratio. When the high-priority queue is empty, the database operation requests in the low-priority queue are processed, and the requests that have exceeded the time slice and have not been processed are moved to the next-level queue, and the lock-free queue concurrency control is completed through atomic operations;

[0144] The size and step size of the performance monitoring sliding window are set, and the mean, variance, coefficient of variation and trend slope of the performance indicators in the sliding window are calculated. When the response time or error rate in the performance indicators exceeds the preset threshold in the system health assessment distribution matrix, the batch size and concurrency in the switching strategy parameter set are adjusted according to the preset ratio until the performance indicators are maintained within the range defined by the system health assessment distribution matrix.

[0145] The integer hash ring is an implementation form that maps data and nodes to a logical ring through an integer hash function, and is used in a consistent hashing algorithm to support fast search and load balancing. The message digest algorithm is an encryption algorithm that generates a fixed-length hash value, which is used to verify data integrity and generate a unique identifier. The skip list is a hierarchical data structure for efficient storage and query of ordered elements. By maintaining skip pointers at different levels, the search complexity is greatly reduced, approaching the efficiency of binary search. The source node is the initial node that executes a request or initiates an operation in a distributed system, and is usually used to mark the starting point of a task.

[0146] Perform risk level assessment. Obtain the affected service node information through the service call chain analysis tool and set a weight value for each node. Taking the e-commerce order system as an example, the order service node weight is 0.9, the payment service node weight is 0.8, and the commodity service node weight is 0.7. Count the service unavailability time after the failure occurs, such as the order service interruption for 15 minutes. At the same time, check the data backup situation. If the most recent backup occurred 30 minutes ago, it is necessary to assess the amount of data that may be lost. Calculate the risk score for the above three indicators according to the weight ratio of 4:3:3. When the risk score exceeds the preset threshold of 0.75, the switching operation is triggered.

[0147] After the switch is triggered, the SHA-256 algorithm is used to calculate the hash value of the physical node. Take a distributed database cluster as an example. There are 5 physical nodes, and each physical node generates 100 virtual nodes. The physical node A is identified as "node-a", and its hash value is located in the range of 0-2^32 of the hash ring. The virtual node identifier is generated by adding a serial number from "-1" to "-100" after the identifier. A skip table is used to store the hash value of the virtual node, and the maximum number of skip tables is set to 16 for fast search of candidate nodes.

[0148] Build a scoring system for the selected candidate nodes. Collect the node's CPU usage, memory usage, network throughput, disk IO and other indicators as status features. Set the experience replay pool capacity to 10,000 records, and randomly select 128 samples for training each time. Calculate the reward value based on the throughput improvement and response time reduction after switching.

[0149] During the decision tree search process, the maximum depth is set to 4, and a maximum of 8 branches are retained in each layer. The node score is composed of the expected performance improvement and the switching risk. 1000 Monte Carlo simulations are performed on each candidate solution, and the average score is taken as the node score. The path with the highest cumulative score is selected as the optimal switching solution.

[0150] The double buffer zero copy mechanism is enabled, and two 32MB buffers are configured for each source node and target node. When the data volume in the first buffer of the source node reaches 24MB, data transmission is triggered. The receiving buffer and the persistent buffer of the target node are the same size, and the buffer switching is controlled by atomic operations.

[0151] Implement multi-level feedback queues and set up 4 priority queues. The time slice of the highest priority queue is 20ms, which decreases to 10ms, 5ms, and 2ms in sequence. Use CAS operations to implement lock-free queue entry and exit. Reduce the priority of timed requests and move them to the next level queue for processing.

[0152] Performance monitoring uses a 60-second sliding window with a step size of 10 seconds. The statistical characteristics of the indicators in the window are calculated. When the response time exceeds 200ms or the error rate exceeds 0.1%, the switching batch size is adjusted to 80% of the original and the concurrency is reduced by 20% until the performance indicators return to normal.

[0153] In this embodiment, the system availability is significantly improved through multi-dimensional risk assessment and intelligent switching mechanism. The node scoring system based on deep policy learning can accurately identify the optimal switching target and reduce the fault recovery time from minutes to seconds. The double-buffered zero-copy mechanism combined with a multi-level feedback queue is used to achieve rapid data migration and request priority scheduling. The adaptive tuning mechanism based on the sliding window realizes dynamic optimization of the switching process. The smoothness of the switching process is ensured through real-time monitoring and automatic parameter adjustment.

[0154] In an optional embodiment,

[0155] If the switching operation is triggered, an integer hash ring is established by a consistent hashing algorithm, a hash value of a physical node identifier is calculated by a message digest algorithm to determine the position of the physical node on the hash ring, an increasing sequence number is added to each physical node identifier to generate a virtual node identifier, a hash value of the virtual node identifier is calculated to determine the position of the virtual node on the hash ring, and the virtual node hash value is stored in a jump table to obtain the candidate node, including:

[0156] Initialize an integer hash ring, wherein the hash ring is stored in a skip table structure, wherein a node of the skip table includes a hash value field and a node information field, wherein the node information field stores a physical node identifier and physical node status information, wherein the physical node status information includes a node load value, an amount of available resources, and a network connection status;

[0157] Read a physical node identifier from a global standby node pool, the physical node identifier including a network address and a port number, calculate a first hash value for the physical node identifier using a message digest algorithm, intercept the first half of the first hash value as a physical node position value, and store the physical node position value and the corresponding physical node identifier in the skip list structure;

[0158] Adding an increasing sequence number after the physical node identifier to generate multiple virtual node identifiers, the increasing sequence number is connected to the physical node identifier through a separator, using the message digest algorithm to calculate the virtual node identifier to obtain a virtual node hash value, intercepting the first half of the virtual node hash value as a virtual node position value, and storing the virtual node position value and the corresponding physical node information in the skip list structure;

[0159] Constructing a multi-level index linked list, wherein the nodes of each level of the index linked list in the multi-level index linked list are arranged in ascending order according to the virtual node position value, and the nodes of the index linked lists at adjacent levels establish jump pointers according to a preset probability, and the jump pointers are used to quickly locate the target interval between the multi-level index linked lists;

[0160] Receive a data migration request, the data migration request including a target data key value, calculate a target hash value of the target data key value using the message digest algorithm, search for a first virtual node greater than or equal to the target hash value starting from the highest-level index linked list in the multi-level index linked list, and use a physical node corresponding to the first virtual node as a target switching node;

[0161] A data transmission channel is established according to the physical node identifier of the target switching node, data on the source node is migrated to the target switching node according to the interval range of the virtual node position value, and the node status information in the jump table structure is updated.

[0162] The virtual node identifier is a unique tag assigned to the virtual node, which is used to implement the identification and distribution of logical nodes in the consistent hashing algorithm. The jump pointer is the core element in the skip list, which is used to connect nodes across levels and support fast positioning of target elements. The highest-level index linked list is the highest-level linked list in the skip list structure, which contains the least elements but can cover the entire data set range, and is used to accelerate top-level searches.

[0163] The data switching method based on the consistent hashing algorithm requires the establishment of an integer hash ring, which is implemented using a skip table structure. Each node in the skip table contains two key fields: the hash value and the node information. The node information field stores the identifier and status information of the physical node, where the status information includes indicators such as the load value of the current node, the number of available resources, and the network connection status.

[0164] During the initialization phase, the system reads the identification information of the physical node from the global standby node pool. The physical node identifier consists of a network address and a port number, such as "192.168.1.100:8080". The system uses a message digest algorithm such as SHA-256 to hash the physical node identifier and obtain a 256-bit hash value. In order to improve search efficiency, the system intercepts the first 128 bits of the hash value as the position value of the physical node on the hash ring, and stores the position value and the node identifier as a key-value pair in the jump table structure.

[0165] Generate multiple virtual nodes for each physical node. Add an increasing serial number after the physical node identifier, and use "#" as a separator between the serial number and the identifier. For example, for the physical node "192.168.1.100:8080", generate virtual node identifiers such as "192.168.1.100:8080#1" and "192.168.1.100:8080#2". The system also uses the SHA-256 algorithm to calculate the hash value of the virtual node identifier and intercepts the first 128 bits as the position value of the virtual node. The position values ​​of these virtual nodes and the corresponding physical node information will be stored in the skip list structure.

[0166] The skip list structure speeds up the search process by building a multi-level index. Each level of the index is an ordered linked list, and the nodes are arranged in ascending order according to the virtual node position value. The index linked lists of adjacent levels are connected by jump pointers, and each node establishes a jump pointer to the previous level with a fixed probability. For example, the probability of establishing a jump pointer upward can be set to 1 / 2, so that on average one node in every two nodes will establish a jump pointer to the previous level.

[0167] When the system receives a data migration request, the request contains the data key value to be migrated. The system uses the same SHA-256 algorithm to calculate the hash value of the data key value, obtains the target hash value, starts searching from the highest-level index of the jump table, and quickly locates the first virtual node that is greater than or equal to the target hash value through the jump pointer. The corresponding physical node is the target node for data migration.

[0168] After determining the target node, the system establishes a data transmission channel based on the identifier of the physical node. The data on the source node is migrated according to the interval range of the virtual node location value, for example, data with hash values ​​in the range of [0x0000, 0xFFFF] is migrated to the target node. After the migration is completed, the status information of the relevant nodes in the jump table is updated, including updating indicators such as node load value and available resources.

[0169] In this embodiment, the node search efficiency is significantly improved by using a skip list structure to store the hash ring. Compared with the traditional sequential search method, the search time complexity is reduced from linear to logarithmic level, which greatly improves the system response speed. The virtual node mechanism is used to effectively solve the problem of uneven data distribution. By generating multiple virtual nodes for each physical node, the data is more evenly distributed among the physical nodes, avoiding the situation where some nodes are overloaded. The hash calculation method based on the message digest algorithm ensures the randomness and uniformity of the data distribution. At the same time, the multi-level index structure of the skip list provides the ability to quickly locate the target interval, which improves the data migration efficiency and load balancing effect of the system as a whole.

[0170] Figure 2 FIG. 1 is a schematic diagram of the structure of the HBase client master-slave switching system based on fault perception according to an embodiment of the present invention. Figure 2 As shown, the system comprises:

[0171] The first unit is used to collect system operation data through the distributed fault monitoring module in the client and form a multi-dimensional system indicator data set through monitoring probes deployed in different network areas, perform data distribution learning on the multi-dimensional system indicator data set through an adaptive clustering algorithm and establish an anomaly detection baseline model, perform a sliding time window operation on the multi-dimensional system indicator data set based on the anomaly detection baseline model and the exponential moving average algorithm in combination with a low-pass filter to obtain pre-processed standard indicator data, perform time-frequency domain feature extraction on the standard indicator data through a multi-scale analysis method based on wavelet packet decomposition, perform dimensionality reduction in combination with a principal component analysis algorithm to obtain a multi-dimensional feature vector, and generate a system health assessment distribution matrix based on a preset deep neural network model and the anomaly detection baseline model;

[0172] The second unit is used to determine the abnormal index based on the system health assessment distribution matrix and compare it with the adaptive threshold calculated by the density clustering algorithm. If it is greater than the adaptive threshold, the fault diagnosis model is triggered, and the fault-related data is collected in a targeted manner through the improved distributed tracing system. The fault-related data is constructed as a service call dependency graph based on the timing graph calculation algorithm combined with the gated graph attention mechanism. The abnormal propagation path is analyzed by a causal reasoning model based on a probability network in combination with the multi-dimensional feature vector, and a fault propagation chain is generated. The system log is obtained, and the system log and the fault propagation chain are associated with the deep semantic analysis model and the graph structure network. The context features are extracted and a fault feature map is generated in combination with the knowledge transfer technology. The fault feature map is multi-layered classified based on the integrated decision tree framework, the features of the abnormal propagation path are integrated, and the fault root cause analysis report and risk level assessment results are output in combination with the hierarchical generalization strategy. The switching strategy parameter set is generated in combination with the dynamic programming algorithm;

[0173] The third unit is used to determine whether to trigger switching based on the risk level assessment result, and if triggered, start the intelligent switching execution module, screen candidate nodes from the global standby node pool through the consistent hashing algorithm combined with the virtual node mechanism and according to the switching policy parameter set, construct a standby node scoring system based on the characteristic indicators in the fault root cause analysis report based on the deep strategy learning algorithm, optimize the decision process through the search tree and perform dynamic scoring, use the candidate node with the highest score value as the switching target node and initialize the double buffer zero copy mechanism based on the fault propagation chain, prioritize database operation requests in combination with the multi-level feedback queue algorithm, perform sequential switching based on the priority division result, compare and model the performance indicators in the switching process with the system health assessment distribution matrix in combination with the timing prediction algorithm, and optimize the switching policy parameter set until the switching is completed.

[0174] According to a third aspect of the embodiments of the present invention,

[0175] An electronic device is provided, comprising:

[0176] processor;

[0177] a memory for storing processor-executable instructions;

[0178] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0179] According to a fourth aspect of the embodiments of the present invention,

[0180] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.

[0181] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.

[0182] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. The HBase client master-slave switching method based on fault perception is characterized in that: include: Collect system operation data through the distributed fault monitoring module in the client and form a multi-dimensional system indicator data set through monitoring probes deployed in different network areas, perform data distribution learning on the multi-dimensional system indicator data set through an adaptive clustering algorithm and establish an anomaly detection baseline model, perform a sliding time window operation on the multi-dimensional system indicator data set based on the anomaly detection baseline model and the exponential moving average algorithm in combination with a low-pass filter to obtain pre-processed standard indicator data, extract time-frequency domain features from the standard indicator data through a multi-scale analysis method based on wavelet packet decomposition, perform dimensionality reduction in combination with a principal component analysis algorithm to obtain a multi-dimensional feature vector, and generate a system health assessment distribution matrix based on a pre-set deep neural network model and the anomaly detection baseline model; Based on the system health assessment distribution matrix, the abnormal index is determined and compared with the adaptive threshold calculated by the density clustering algorithm. If it is greater than the adaptive threshold, the fault diagnosis model is triggered, and the fault-related data is collected in a targeted manner through the improved distributed tracing system. The fault-related data is constructed as a service call dependency graph based on the timing graph calculation algorithm combined with the gated graph attention mechanism. The abnormal propagation path is analyzed by a causal reasoning model based on a probability network in combination with the multi-dimensional feature vector, and a fault propagation chain is generated. The system log is obtained, and the system log and the fault propagation chain are associated with each other based on a deep semantic analysis model and a graph structure network. The context features are extracted and a fault feature map is generated in combination with the knowledge transfer technology. The fault feature map is multi-layered classified based on the integrated decision tree framework, and the features of the abnormal propagation path are integrated. The fault root cause analysis report and the risk level assessment result are output in combination with the hierarchical generalization strategy, and the switching strategy parameter set is generated in combination with the dynamic programming algorithm. Based on the risk level assessment result, it is determined whether to trigger switching. If triggered, the intelligent switching execution module is started, and candidate nodes are screened from the global standby node pool through the consistent hashing algorithm combined with the virtual node mechanism and according to the switching policy parameter set. Based on the deep policy learning algorithm, the characteristic indicators in the fault root cause analysis report are used to construct a standby node scoring system. The decision process is optimized through the search tree and dynamic scoring is performed. The candidate node with the highest score value is used as the switching target node and the double buffer zero copy mechanism is initialized based on the fault propagation chain. The database operation request is prioritized in combination with the multi-level feedback queue algorithm, and sequential switching is performed based on the priority division result. The performance indicators in the switching process are compared with the system health assessment distribution matrix in combination with the timing prediction algorithm to model, and the switching policy parameter set is optimized until the switching is completed.

2. The method according to claim 1, characterized in that The distributed fault monitoring module in the client collects system operation data and forms a multi-dimensional system indicator data set through monitoring probes deployed in different network areas. The adaptive clustering algorithm is used to learn the data distribution of the multi-dimensional system indicator data set and establish an anomaly detection baseline model, including: The distributed fault monitoring module collects system operation data, and the system operation data includes processor utilization index, memory occupancy index, disk performance index, and network performance index. The processor utilization index is obtained by collecting the user state utilization rate, system state utilization rate, input and output waiting rate, and idle rate of each processor core. The memory occupancy index is obtained by collecting the physical memory usage, virtual memory usage, cache usage, and buffer usage. The disk performance index is obtained by collecting the read and write rate, average response time, input and output queue length, and device utilization of each disk device. The network performance index is obtained by collecting the input and output bandwidth utilization, network delay, packet loss rate, and retransmission rate. The distributed fault monitoring module collects the system operation data through monitoring probes deployed in different network areas to form a multi-dimensional system indicator data set, wherein the monitoring probes are composed of a first-level probe node, a second-level probe node, and a third-level probe node. The first-level probe node is configured according to the ratio of the number of racks and is responsible for local data collection. The second-level probe node is configured according to the number of computer rooms and is responsible for aggregating the data collected by the first-level probe node for preprocessing. The third-level probe node is configured after dividing the area according to the geographical location and network topology structure and is responsible for coordinating and managing the data collection tasks in the area. The first-level probe node, the second-level probe node, and the third-level probe node maintain connection through a heartbeat mechanism. After the multidimensional system indicator data set is standardized and converted, the Euclidean distance between the sample and the cluster center is calculated based on the initial cluster center for classification. When the variance of the samples in the cluster exceeds the preset threshold, a new cluster center is added in the direction of the maximum variance in the current cluster. When the distance between adjacent cluster centers is less than the preset threshold, adjacent clusters are merged. If the moving distance of the cluster center is less than the preset convergence threshold for 100 consecutive iterations, the data distribution learning is considered to be completed. Based on the results of data distribution learning, the mean and standard deviation of samples in each cluster are calculated, and the multiple range of the mean and the standard deviation is set as the normal value range of the current indicator. At the same time, the historical data is divided into hourly cycles to establish an hourly baseline, a daily baseline is established according to a daily cycle, and a monthly baseline is established according to a monthly cycle. The hourly baseline, the daily baseline, and the monthly baseline are weighted to form the anomaly detection baseline model.

3. The method according to claim 2, characterized in that Based on the anomaly detection baseline model and the exponential moving average algorithm, a sliding time window operation is performed on the multidimensional system indicator data set in combination with a low-pass filter to obtain preprocessed standard indicator data, time-frequency domain feature extraction is performed on the standard indicator data through a multi-scale analysis method based on wavelet packet decomposition, and a multi-dimensional feature vector is obtained by dimensionality reduction in combination with a principal component analysis algorithm, and a system health assessment distribution matrix is ​​generated based on a pre-set deep neural network model and the anomaly detection baseline model, including: The distributed fault monitoring module performs a sliding time window operation on the multidimensional system indicator data set based on the anomaly detection baseline model and the exponential moving average algorithm in combination with a low-pass filter, fills the continuous indicator with linear interpolation of the two normal values ​​before and after, and fills the discrete indicator with the nearest neighbor normal value, distinguishes high-frequency change indicators from low-frequency change indicators based on the ratio of the standard deviation to the mean, and smoothes the high-frequency change indicators and the low-frequency change indicators with different smoothing factors to obtain pre-processed standard indicator data; A multi-scale analysis method based on wavelet packet decomposition is used to extract time-frequency domain features of the standard indicator data, decompose the standard indicator data into multiple frequency bands, calculate the energy value of each frequency band, eliminate frequency bands whose energy ratio is lower than a preset energy ratio threshold, extract the mean, variance, skewness, kurtosis, maximum value and minimum value of the retained frequency bands as statistical features, and calculate the energy ratio between different frequency bands based on the statistical features to obtain time-frequency domain features; The distributed fault monitoring module calculates the covariance matrix of the feature matrix based on the principal component analysis algorithm to obtain eigenvalues ​​and eigenvectors, selects eigenvectors corresponding to the eigenvalues ​​based on a preset variance contribution rate threshold to construct a conversion matrix, and maps the time-frequency domain features to the principal component space to obtain a multi-dimensional eigenvector; The distributed fault monitoring module inputs the multi-dimensional feature vector into a preset deep neural network model, wherein the number of input layer nodes of the deep neural network model is the same as the number of principal components, the number of hidden layer nodes is set according to the ratio of the input dimensions, the number of output layer nodes is the same as the type of monitoring indicators, the parameters of the deep neural network model are trained based on normal state data, the multi-dimensional feature vector is input into the trained deep neural network model to obtain a health score, and normalized in combination with the anomaly detection baseline model to generate a system health assessment distribution matrix.

4. The method according to claim 1, characterized in that: Based on the system health assessment distribution matrix, the abnormal index is determined and compared with the adaptive threshold calculated by the density clustering algorithm. If it is greater than the adaptive threshold, the fault diagnosis model is triggered. The fault-related data is collected in a targeted manner through the improved distributed tracing system. The fault-related data is constructed as a service call dependency graph based on the timing graph calculation algorithm combined with the gated graph attention mechanism. The abnormal propagation path is analyzed by a causal reasoning model based on a probability network in combination with the multi-dimensional feature vector, and a fault propagation chain is generated. The system log is obtained, and the system log and the fault propagation chain are associated with the deep semantic analysis model and the graph structure network. The context features are extracted and the fault feature map is generated in combination with the knowledge transfer technology. The fault feature map is multi-layered classified based on the integrated decision tree framework, the features of the abnormal propagation path are integrated, and the fault root cause analysis report and the risk level assessment result are output in combination with the hierarchical generalization strategy. The switching strategy parameter set is generated in combination with the dynamic programming algorithm, including: The health score of each indicator in the system health assessment distribution matrix is ​​compared with the historical health data, the historical health data is grouped according to the indicator dimension, the distance between each data point and the adjacent data point is calculated to obtain the local data density, the local data density is sorted and the difference between the adjacent density values ​​is calculated, the density demarcation point is determined at the position where the difference jumps, the data is divided into multiple density intervals based on the density demarcation point, the highest density point is selected as the density center point in each density interval, the average distance from all data points in the density interval to the density center point is calculated as the characteristic radius of the current density interval, and the adaptive threshold is determined by the density weighting method of the density interval; When the health score is lower than the adaptive threshold, a service call topology is constructed through a distributed tracing system, a tracing identifier is injected at the abnormal service node, the tracing identifier is transmitted along with the request on the call link, call data information is recorded for the request injected with the tracing identifier, full sampling is adopted for the node marked abnormal, and a sampling rate of 50% is adopted for the associated upstream and downstream nodes; Aggregate call data with the same tracking identifier and construct a call sequence according to timestamps, perform semantic analysis on interface names in the call chain to calculate the similarity between interfaces, classify interfaces with high similarity into the same call group, analyze the call association features between the call groups, assign importance weights to each call group based on the call association features, and construct a weighted service call dependency graph based on the association degree between call groups and the importance weights; The abnormal probability of the adjacent nodes when each service node in the service call dependency graph is abnormal is counted to establish a conditional probability relationship, and the abnormal state propagation sequence is analyzed in combination with the abnormal occurrence sequence of the indicators recorded in the multi-dimensional feature vector. The diffusion probability distribution of the abnormal event in the network is obtained through iterative calculation, and the abnormal propagation path is traced back to locate the abnormal source node. The main abnormal propagation path is screened based on the probability dependence strength between nodes to form a fault propagation chain; Extract structured information fields from system logs, identify key information in log texts through word frequency analysis to build log templates, map log texts to the log templates and extract variable values ​​to capture changes in system operation status, perform vectorization on log texts to extract semantic features, convert topological structure information of the fault propagation chain into feature vectors, and fuse the semantic features with the feature vectors through an attention mechanism to form a fault context feature vector; Retrieve cases with similar symptom characteristics to the current fault from the historical fault case library to extract fault handling experience, migrate historical fault modes to the current scenario through feature distribution adaptation, build a fault feature map based on the current fault manifestation, use a three-layer classification structure to locate the fault, perform coarse-grained classification based on symptom characteristics, refine the classification results based on the abnormal propagation path, and determine the fault type by integrating system status characteristics. Use the random forest method to integrate multiple classifier results at each layer of classification. Based on the fault type, the current state of the system is determined and the system state vector is constructed. The impact of the switching operation on the system state is analyzed, the degree of system impact during the switching process is evaluated, and the switching path with the minimum system loss is obtained through iterative calculation of the state value, and a system switching strategy parameter set is generated.

5. The method according to claim 4, characterized in that Retrieve cases similar to the current fault symptom characteristics from the historical fault case library to extract fault handling experience, migrate historical fault modes to the current scenario through feature distribution adaptation, build a fault feature map based on the current fault manifestation, and locate the fault using a three-layer classification structure. Perform coarse-grained classification based on symptom characteristics, refine the classification results based on the abnormal propagation path, and determine the fault type by integrating system status characteristics. Each layer of classification uses the random forest method to integrate multiple classifiers. The results include: Extract key features of historical fault cases, including system layer indicators, service layer indicators, component layer indicators and application layer indicators, segment the key features into time windows, the time window length is five minutes, the time window overlap time is one minute, calculate statistical feature values ​​within the time window, the statistical feature values ​​include mean, standard deviation, maximum value, minimum value, median, skewness, kurtosis and quantile, retain the original value for continuous features with a fixed value range, perform maximum and minimum value normalization on continuous features with inconsistent dimensions, and perform one-hot encoding conversion on discrete features; Collect the monitoring data of the current fault in the last thirty minutes, extract features from the monitoring data to obtain the current fault vector, calculate the cosine similarity between the current fault vector and the historical fault case vector, select the historical fault cases with a cosine similarity greater than 0.8 as the candidate set, and reduce the cosine similarity threshold to 0.6 when the number of cases in the candidate set is less than ten; Extracting environmental context information, the environmental context information includes hardware configuration information, deployment information and load feature information, constructing a feature graph, taking the key features as graph nodes, calculating the Pearson correlation coefficient between the graph nodes, establishing edges between the graph nodes when the absolute value of the Pearson correlation coefficient is greater than 0.6, taking the Pearson correlation coefficient as the edge weight, grouping the graph nodes using a hierarchical clustering method, setting a minimum distance threshold of 0.4, and using the graph edit distance to calculate the pattern similarity between the result of the hierarchical clustering and the combination of historical fault features in the candidate set; A multi-layer random forest classifier is used for fault classification. The first-layer classifier divides faults into major categories based on system resource indicator subsets, service status indicator subsets, error log feature subsets, alarm information subsets and performance indicator subsets. The second-layer classifier subdivides faults into subcategories based on the timing relationship of abnormal nodes, propagation delay time, number of affected services and similarity of error types. The third-layer classifier determines the specific cause of the fault based on the system operation status curve, load change trend and resource usage fluctuation. The voting weight is determined according to the accuracy of each layer of classifiers on the validation set, and the predicted fault type is obtained through weighted voting.

6. The method according to claim 1, characterized in that Based on the risk level assessment result, it is determined whether to trigger the switch. If triggered, the intelligent switch execution module is started. The candidate nodes are screened from the global standby node pool through the consistent hashing algorithm combined with the virtual node mechanism and according to the switch policy parameter set. The standby node scoring system is constructed based on the characteristic indicators in the fault root cause analysis report based on the deep strategy learning algorithm. The decision process is optimized through the search tree and dynamic scoring is performed. The candidate node with the highest score value is used as the switching target node and the double buffer zero copy mechanism is initialized based on the fault propagation chain. The database operation request is prioritized in combination with the multi-level feedback queue algorithm. Sequential switching is performed based on the priority division result. The performance indicators in the switching process are compared and modeled with the system health assessment distribution matrix in combination with the timing prediction algorithm. The switching policy parameter set is optimized until the switching is completed, including: Obtain the risk level assessment result, determine the number of affected service nodes in combination with the service call chain, set a weight value according to the importance of the service node in the business system, and the weight value ranges from zero to one. Count the service unavailability time caused by the fault, count the amount of new data after the most recent data backup, calculate the impact score of the service node, the service unavailability time score, and the data loss risk score according to the preset weight ratio to obtain a fault risk score, and determine whether to trigger the switching operation based on the fault risk score; If the switching operation is triggered, an integer hash ring is established through a consistent hashing algorithm, a hash value of the physical node identifier is calculated using a message digest algorithm to determine the position of the physical node on the hash ring, an increasing serial number is appended to each physical node identifier to generate a virtual node identifier, a hash value of the virtual node identifier is calculated to determine the position of the virtual node on the hash ring, and the virtual node hash value is stored in a jump list to obtain the candidate node; Based on the candidate nodes, a state feature vector including current load indicators, resource usage, and network status is constructed through a deep strategy learning algorithm, the selection combination of the candidate nodes is used as the action space, the reward value is calculated based on the degree of improvement in system performance after switching, a sample pool of preset capacity is maintained from historical switching records, and training samples are randomly extracted for batch update; Initialize the decision tree and record the candidate nodes, performance indicator values, number of alternatives and cumulative path scores at the decision tree nodes, select the candidate with the highest expected score for node expansion based on the scoring system, perform random sampling simulation on each layer of the decision tree and calculate the average score as the node score, and determine the optimal switching path according to the preset depth and branch number limit; Based on the optimal switching path, a double buffer zero copy mechanism is started, two buffers of the same size are respectively set at the source node and the switching target node, the source node writes into the first buffer and transmits the second buffer data to the switching target node through direct memory access, when the first buffer is full, the data of the first buffer is switched to the second buffer to be written and transmitted, and the switching target node uses the receiving buffer and the persistent buffer to process the transmitted data alternately; Based on the double-buffered zero-copy mechanism, multiple priority queues are constructed using a ring buffer, and the time slice size of each priority queue is set according to a preset ratio. When the high-priority queue is empty, the database operation requests in the low-priority queue are processed, and the requests that have exceeded the time slice and have not been processed are moved to the next-level queue, and the lock-free queue concurrency control is completed through atomic operations; The size and step size of the performance monitoring sliding window are set, and the mean, variance, coefficient of variation and trend slope of the performance indicators in the sliding window are calculated. When the response time or error rate in the performance indicators exceeds the preset threshold in the system health assessment distribution matrix, the batch size and concurrency in the switching strategy parameter set are adjusted according to the preset ratio until the performance indicators are maintained within the range defined by the system health assessment distribution matrix.

7. The method according to claim 6, characterized in that If the switching operation is triggered, an integer hash ring is established by a consistent hashing algorithm, a hash value of a physical node identifier is calculated by a message digest algorithm to determine the position of the physical node on the hash ring, an increasing sequence number is added to each physical node identifier to generate a virtual node identifier, a hash value of the virtual node identifier is calculated to determine the position of the virtual node on the hash ring, and the virtual node hash value is stored in a jump table to obtain the candidate node, including: Initialize an integer hash ring, wherein the hash ring is stored in a skip table structure, wherein a node of the skip table includes a hash value field and a node information field, wherein the node information field stores a physical node identifier and physical node status information, wherein the physical node status information includes a node load value, an amount of available resources, and a network connection status; Read a physical node identifier from a global standby node pool, the physical node identifier including a network address and a port number, calculate a first hash value for the physical node identifier using a message digest algorithm, intercept the first half of the first hash value as a physical node position value, and store the physical node position value and the corresponding physical node identifier in the skip list structure; Adding an increasing sequence number after the physical node identifier to generate multiple virtual node identifiers, the increasing sequence number is connected to the physical node identifier through a separator, using the message digest algorithm to calculate the virtual node identifier to obtain a virtual node hash value, intercepting the first half of the virtual node hash value as a virtual node position value, and storing the virtual node position value and the corresponding physical node information in the skip list structure; Constructing a multi-level index linked list, wherein the nodes of each level of the index linked list in the multi-level index linked list are arranged in ascending order according to the virtual node position value, and the nodes of the index linked lists at adjacent levels establish jump pointers according to a preset probability, and the jump pointers are used to quickly locate the target interval between the multi-level index linked lists; Receive a data migration request, the data migration request including a target data key value, calculate a target hash value of the target data key value using the message digest algorithm, search for a first virtual node greater than or equal to the target hash value starting from the highest-level index linked list in the multi-level index linked list, and use a physical node corresponding to the first virtual node as a target switching node; A data transmission channel is established according to the physical node identifier of the target switching node, data on the source node is migrated to the target switching node according to the interval range of the virtual node position value, and the node status information in the jump table structure is updated.

8. An HBase client master-slave switching system based on fault perception, used to implement the method described in any one of claims 1 to 7, characterized in that: include: The first unit is used to collect system operation data through the distributed fault monitoring module in the client and form a multi-dimensional system indicator data set through monitoring probes deployed in different network areas, perform data distribution learning on the multi-dimensional system indicator data set through an adaptive clustering algorithm and establish an anomaly detection baseline model, perform a sliding time window operation on the multi-dimensional system indicator data set based on the anomaly detection baseline model and the exponential moving average algorithm in combination with a low-pass filter to obtain pre-processed standard indicator data, perform time-frequency domain feature extraction on the standard indicator data through a multi-scale analysis method based on wavelet packet decomposition, perform dimensionality reduction in combination with a principal component analysis algorithm to obtain a multi-dimensional feature vector, and generate a system health assessment distribution matrix based on a preset deep neural network model and the anomaly detection baseline model; The second unit is used to determine the abnormal index based on the system health assessment distribution matrix and compare it with the adaptive threshold calculated by the density clustering algorithm. If it is greater than the adaptive threshold, the fault diagnosis model is triggered, and the fault-related data is collected in a targeted manner through the improved distributed tracing system. The fault-related data is constructed as a service call dependency graph based on the timing graph calculation algorithm combined with the gated graph attention mechanism. The abnormal propagation path is analyzed by a causal reasoning model based on a probability network in combination with the multi-dimensional feature vector, and a fault propagation chain is generated. The system log is obtained, and the system log and the fault propagation chain are associated with the deep semantic analysis model and the graph structure network. The context features are extracted and a fault feature map is generated in combination with the knowledge transfer technology. The fault feature map is multi-layered classified based on the integrated decision tree framework, the features of the abnormal propagation path are integrated, and the fault root cause analysis report and risk level assessment results are output in combination with the hierarchical generalization strategy. The switching strategy parameter set is generated in combination with the dynamic programming algorithm; The third unit is used to determine whether to trigger switching based on the risk level assessment result, and if triggered, start the intelligent switching execution module, screen candidate nodes from the global standby node pool through the consistent hashing algorithm combined with the virtual node mechanism and according to the switching policy parameter set, construct a standby node scoring system based on the characteristic indicators in the fault root cause analysis report based on the deep strategy learning algorithm, optimize the decision process through the search tree and perform dynamic scoring, use the candidate node with the highest score value as the switching target node and initialize the double buffer zero copy mechanism based on the fault propagation chain, prioritize database operation requests in combination with the multi-level feedback queue algorithm, perform sequential switching based on the priority division result, compare and model the performance indicators in the switching process with the system health assessment distribution matrix in combination with the timing prediction algorithm, and optimize the switching policy parameter set until the switching is completed.

9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Perceptual information integrated access system based on urban important infrastructure

    CN113642946A

  • Fault-tolerant method for improving underwater robot networking robustness

    CN118741573A