Deep learning-based unified information UOS system fault rapid repair method

Through deep learning technology, the faults of the Tongxin UOS system are handled, the fault types are identified using feature engineering and BiLSTM-multi-head attention model, and the repair scheme is verified in the sandbox environment, which solves the problem of inefficient failure handling in the existing technology and achieves fast and safe fault repair.

CN120523641AInactive Publication Date: 2025-08-22WUHAN DEFA ELECTRONIC INFORMATION CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511025678.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-24
Publication Date
2025-08-22
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing Tongxin UOS system fault repair methods rely too much on the preset rule library and cannot handle new unknown faults. The secondary risks brought by repair cannot be controlled, and manual system fault detection is inefficient.

Method used

The rapid fault repair method based on deep learning is used to obtain the initial system status index data through system commands and data acquisition tools, perform feature engineering processing, and use BiLSTM-DNN fault classifier model for multi-head attention to determine the fault type, and verify the feasibility of the repair solution in a sandbox environment to generate a fault repair solution.

Benefits of technology

Effectively identify new unknown faults, quickly generate repair solutions that take into account efficiency and resource utilization, improve the efficiency, accuracy and safety of system fault handling, avoid repair risks, and ensure the stable and reliable operation of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120523641A_ABST
    Figure CN120523641A_ABST
Patent Text Reader

Abstract

The invention discloses a deep learning-based unified information UOS system fault rapid repair method, and relates to the technical field of UOS system repair. The method comprises the following steps: acquiring initial system state index data based on a system command and a data acquisition tool; performing feature engineering processing on the obtained initial system state index data to obtain a unified information DOS system feature vector; determining whether the unified information DOS system has a fault and the fault type based on the unified information DOS system feature vector, and obtaining a fault recovery scheme after determining that the unified information DOS system has the fault; the feasibility and repair risk of a fault repair scheme are verified based on a sandbox environment, if the fault repair scheme is feasible, the scheme is executed, and if the fault repair scheme is not feasible, the repair scheme is regenerated, so that the problems that a traditional operating system repair tool excessively depends on a preset rule base, novel unknown faults cannot be processed, secondary risks caused by repair cannot be controlled, and the repair efficiency is low are solved. The efficiency of manual system fault checking is low, and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of UOS system repair technology, and specifically to a method for quickly repairing Tongxin UOS system faults based on deep learning. Background Art

[0002] Tongxin UOS is a full-featured operating system that can be used in a variety of scenarios. In terms of office, it can support various office software, allowing users to efficiently process documents, spreadsheets, presentations, etc.; in daily use, it can meet entertainment needs such as browsing the web, watching videos, listening to music, and can also perform file management, system settings and other operations. When processing tasks, Tongxin UOS uses its kernel scheduling resources to reasonably allocate CPU, memory and other hardware resources to various applications to ensure stable operation and efficient response of the system. At the same time, it uses a graphical interface to provide users with intuitive and convenient operation methods. Users interact with the system through mouse clicks, keyboard input, etc. to complete various commands, such as starting applications, opening files, adjusting system parameters, etc., and can also perform more advanced system management and operations through the command line interface.

[0003] Tongxin UOS is widely used in various fields, including government, enterprise, education, healthcare, and finance. In the government sector, it supports application scenarios such as high-availability clusters, middleware, and cloud computing to meet the needs of government information construction. In the enterprise sector, it provides standardized services, virtualization, cloud computing support, and other functions to facilitate the digital transformation of enterprises. In the education sector, it covers multiple scenarios such as teaching management to meet the needs of users such as education administrators, teachers, and students.

[0004] Tongxin UOS offers a variety of troubleshooting methods. When encountering a system software failure, you can repair or reinstall the damaged software package using the system's built-in package management tools. For some system configuration issues, you can enter the system's safe mode and attempt to restore the default configuration or adjust relevant parameters. If problems such as hardware driver incompatibility arise, you can download the latest adapter driver from the Tongxin official website and install it. However, existing Tongxin UOS troubleshooting methods suffer from over-reliance on a pre-set rule base, an inability to handle new, unknown faults, uncontrollable secondary risks associated with repairs, and inefficient manual troubleshooting of system faults. Summary of the Invention

[0005] In response to the shortcomings of the existing technology, the present invention provides a deep learning-based method for quickly repairing Tongxin UOS system faults, which solves the problems of existing traditional system repair tools that are overly dependent on preset rule bases, cannot handle new unknown faults, cannot control the secondary risks brought about by repairs, and are inefficient in manual troubleshooting of system faults.

[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: a deep learning-based Tongxin UOS system fault rapid repair method, comprising the following steps: obtaining initial system status indicator data based on system commands and data acquisition tools; performing feature engineering processing on the obtained initial system status indicator data to obtain the Tongxin UOS system feature vector; determining whether there is a fault in the Tongxin UOS system and the type of fault based on the Tongxin UOS system feature vector, and obtaining a fault repair plan after determining that there is a fault in the Tongxin UOS system; verifying the fault repair plan based on a sandbox environment, and if the fault repair plan is feasible, executing the plan; if it is not feasible, regenerating the repair plan.

[0007] The Tongxin UOS system feature vector is obtained by performing feature engineering processing on the obtained initial system status indicator data, which includes the following steps: after data cleaning, the initial system status indicator data is segmented based on the word segmentation tool to obtain status indicator words; the word frequency and inverse document frequency of each status indicator word are calculated, and the word frequency and inverse document frequency are weighted and summed to obtain the TF-IDF value of each status indicator word; each status indicator word is sorted in descending order according to the TF-IDF value to obtain the initial Tongxin UOS system feature vector; the initial Tongxin UOS system feature vector is dimensionally adjusted to obtain the Tongxin UOS system feature vector.

[0008] The dimension of the initial UOS system feature vector is adjusted to obtain the UOS system feature vector, including the following steps: judging whether the feature vector dimension of the initial Tongxin UOS system feature vector is greater than the set K value: if the feature vector dimension is greater than the set K value, performing principal component analysis on each feature in the initial UOS system feature vector, re-sorting the state indicator words corresponding to each feature according to the variance contribution, and obtaining the first K state indicator words to construct the UOS system feature vector; if the feature vector dimension is not greater than the set K value, performing zero-value padding to obtain a UOS system feature vector with a feature vector dimension of K.

[0009] A fault repair plan is further obtained, including the following steps: inputting the Tongxin UOS system feature vector into a trained fault type diagnosis model to determine whether the Tongxin UOS system has a fault and the fault type; if the Tongxin UOS system has a fault, obtaining the fault state vector of the Tongxin UOS system according to the determined fault type; and obtaining a fault repair plan according to the fault state vector of the Tongxin UOS system.

[0010] The fault type diagnosis model is a BiLSTM-multi-head attention DNN fault classifier model, in which the BiLSTM has 3 layers and the multi-head attention has 4 heads: the first layer of the BiLSTM includes 64 neurons, the second layer of the BiLSTM includes 128 neurons, and the third layer of the BiLSTM includes 64 neurons.

[0011] Obtaining a fault state vector of a Tongxin UOS system and obtaining a fault repair plan according to the fault state vector comprises the following steps: determining a state vector set for representing the fault state of the Tongxin UOS system based on a determined fault type and a fault type-state feature combination mapping table stored in a database; obtaining corresponding features from initial system state indicator data based on the state vector set, and arranging them into a fault state vector of the Tongxin UOS system in a set order; determining a repair plan set corresponding to a current fault state vector based on the determined fault state vector and a fault state vector-repair plan set mapping table stored in a database; evaluating each initial repair plan in the repair plan set using a reinforcement learning reward function to obtain a repair action evaluation value corresponding to each initial repair plan; and defining the initial repair plan corresponding to the maximum repair action evaluation value as the fault repair plan.

[0012] The reinforcement learning reward function is expressed as: ; in, is the output value of the reinforcement learning reward function, is the probability that the initial repair solution a successfully repairs the fault under the current fault state vector s, is the time saved by the initial repair solution a to successfully repair the fault under the current fault state vector s, is the resource occupancy rate caused by the initial repair solution a successfully repairing the fault under the current fault state vector s, for The weight factor, for The weight factor, for The weight factor of .

[0013] Verifying the fault repair plan based on the sandbox environment includes the following steps: creating a sandbox environment in the Tongxin UOS system that is completely consistent with the Tongxin UOS system, and restoring the system state of the sandbox environment to the same state as before the Tongxin UOS system fault occurred; following the operating steps of the fault repair plan, executing repair commands or configuration modifications in the sandbox environment in sequence to obtain multi-dimensional evaluation indicators, including repair success rate, execution time efficiency and rollback feasibility; obtaining a comprehensive score based on the multi-dimensional evaluation indicators, and verifying the feasibility of the fault repair plan based on the comprehensive score.

[0014] The feasibility of the fault repair plan is verified based on the comprehensive score, including the following steps: obtaining a set threshold stored in a database; if the comprehensive score is greater than or equal to the set threshold and there is no serious system-level risk, the fault repair plan is feasible; if the comprehensive score is less than the set threshold or there is a serious system-level risk, the fault repair plan is not feasible.

[0015] The method for obtaining the comprehensive score is as follows: ; in, For the comprehensive score, For the repair success rate, For execution time efficiency, is the actual execution time, is the mean repair time threshold stored in the database, For rollback feasibility, for The weight factor, for The weight factor, for The weight factor of .

[0016] The present invention has the following beneficial effects: This deep learning-based rapid fault repair method for the Tongxin UOS system first uses system commands and collection tools to obtain initial system status indicator data, and then obtains system feature vectors through feature engineering processing to determine the fault and type and generate a repair plan. Finally, the feasibility and risk of the plan are verified through a sandbox environment. It gets rid of the dependence on the preset rule base, can effectively identify new unknown faults, and quickly generate repair plans that take into account efficiency and resource utilization. With the help of the sandbox environment, the system security is guaranteed and the repair risks are avoided. It comprehensively improves the efficiency, accuracy and security of system fault handling, and effectively ensures the stable and reliable operation of the Tongxin UOS system.

[0017] Of course, any product implementing the present invention does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 This is a flow chart of the Tongxin UOS system fault rapid repair method based on deep learning of the present invention. DETAILED DESCRIPTION

[0019] See also Figure 1 , an embodiment of the present invention provides a technical solution: a method for quickly repairing a Tongxin UOS system fault based on deep learning, comprising the following steps: obtaining initial system status indicator data based on system commands and data acquisition tools; Initial system status indicator data covers multiple system metrics, such as the CPU user-mode / kernel-mode time ratio, memory leak detection counters, file system inode error rates, and network TCP retransmission rates. These data are obtained using system commands and tools, including system logs obtained using journalctl -p 3 -b, process information obtained using psutil.process_iter(), disk information obtained using df -h, and network information obtained using ip a. This data forms the basis for subsequent analysis and reflects the current system status.

[0020] The initial system status indicator data is processed through feature engineering to obtain the Tongxin UOS system feature vector. This is used to determine whether the Tongxin UOS system has any faults and the specific fault type. Fault types include hardware faults (such as ACPI exceptions and DMA failures), software faults (such as service startup timeouts and dependency loop detection), and network faults (such as DNS pollution and IPsec negotiation failures).

[0021] After data cleaning of the initial system status indicator data, word segmentation is performed based on the word segmentation tool to obtain status indicator words; the word frequency and inverse document frequency of each status indicator word are calculated, and the word frequency and inverse document frequency are weighted and summed to obtain the TF-IDF value of each status indicator word; each status indicator word is sorted in descending order according to the TF-IDF value to obtain the initial Tongxin UOS system feature vector; the initial Tongxin UOS system feature vector is dimensionally adjusted to obtain the Tongxin UOS system feature vector.

[0022] Clean the collected initial system status indicator data to remove special characters, punctuation, and irrelevant spaces. Special characters and punctuation may appear in logs as errors, but they are not useful for fault feature analysis. Removing them makes the text more concise and standardized. Segment the cleaned text to break long text into individual words or phrases. For Chinese logs, use professional word segmentation tools for accurate segmentation; for English logs, segment by spaces.

[0023] Term frequency refers to how often a term appears in a single log entry. A higher term frequency indicates a more important term in that log entry. It is calculated by dividing the number of times a term appears in a log entry by the total number of terms in that log entry. Inverse document frequency (IDF) reflects the general importance of a term within the entire log dataset. If a term appears frequently in most log entries, its inverse document frequency (IDF) is low; conversely, if it appears in only a few log entries, its inverse document frequency (IDF) is high. To calculate the inverse document frequency, first count the number of log entries containing the term, then divide the total number of log entries by the number of log entries containing the term, and then take the logarithm. Multiplying the term frequency by the inverse document frequency (IDF) yields the TF-IDF value for each term. The TF-IDF value comprehensively considers the importance of a term within a single log entry and its general importance within the entire log dataset. A higher value indicates that a term is more critical for distinguishing different log types and identifying fault characteristics.

[0024] All words are sorted by TF-IDF value, and words with high TF-IDF values ​​are selected as features. These words best represent the key information in the log text, helping to improve the accuracy of fault diagnosis. The selected features are arranged in a specific order to construct a vector space. For each log text, the TF-IDF value of each feature word is calculated and these values ​​are sequentially assigned to the corresponding positions in the vector space to obtain the feature vector for the log text.

[0025] Determine whether the eigenvector dimension of the initial Tongxin UOS system eigenvector is greater than the set K value: if the eigenvector dimension is greater than the set K value, perform principal component analysis on each feature in the initial UOS system eigenvector, reorder the state indicator words corresponding to each feature according to the variance contribution, and obtain the first K state indicator words to construct the UOS system eigenvector; if the eigenvector dimension is not greater than the set K value, perform zero-value padding to obtain a UOS system eigenvector with a eigenvector dimension of K.

[0026] If K is 128, then the dimension of the initial Tongxin UOS system feature vector should be 128. If the dimension of the constructed feature vector is higher than 128, it is necessary to use dimensionality reduction techniques such as principal component analysis (PCA) to reduce the vector dimension while retaining key information. If the dimension is lower than 128, it can be supplemented to 128 dimensions by filling with zero values ​​or using feature expansion methods.

[0027] Instead of manually defining "which features are fault signatures," a data-driven approach mines key features from all possible terms. Even if an unknown fault occurs, the model can detect it as long as the associated terms in the log have a unique TF-IDF distribution (e.g., low document frequency but high word frequency). Using scripts or tools to batch process system command output (such as journalctl logs and psutil process information), data cleaning, word segmentation, and TF-IDF calculation are completed in seconds, improving processing efficiency by hundreds of times. This ensures that all potential fault-related terms are evaluated equally, reducing the risk of missed detections.

[0028] Based on the Tongxin UOS system feature vector, the system determines whether a fault exists in the Tongxin UOS system and the type of fault. Once a fault is confirmed, a fault repair plan is generated. The system feature vector is input into the fault type diagnosis model, which quickly outputs the presence and type of fault in the system, significantly reducing fault diagnosis time compared to manual troubleshooting. A repair plan is automatically generated based on the fault state vector, eliminating the tedious process of manually searching for repair methods. A reinforcement learning reward function is used to evaluate the solutions in the repair plan set and select the optimal solution, reducing the uncertainty and trial-and-error costs associated with manual repair plan selection.

[0029] The Tongxin UOS system feature vector is input into a trained fault type diagnosis model to determine whether the Tongxin UOS system is faulty and what the fault type is. If a fault is present, the fault state vector of the Tongxin UOS system is derived based on the determined fault type. A repair plan is then derived from the fault state vector. Automatically extracting the state vector and a set of repair plans from a mapping table based on the fault type significantly reduces the time required to manually search and match repair plans. A reinforcement learning reward function quantitatively evaluates the repair plans and automatically selects the one with the highest score as the fault repair plan. Compared to subjective judgment and manual selection of repair plans, this approach is more scientific and efficient, avoiding the need for trial and error.

[0030] The fault type diagnosis model is a BiLSTM-multi-head attention DNN fault classifier model, in which the BiLSTM has 3 layers and the multi-head attention has 4 heads: the first layer of the BiLSTM includes 64 neurons, the second layer of the BiLSTM includes 128 neurons, and the third layer of the BiLSTM includes 64 neurons.

[0031] The BiLSTM network is capable of capturing time series information within feature vectors, as system operational data often exhibits temporal relationships. For example, certain faults may develop gradually over time. The first BiLSTM layer, consisting of 64 neurons, performs preliminary processing on the input feature vector, extracting essential feature information. These neurons utilize gating mechanisms (input gate, forget gate, and output gate) to control the flow and memory of information, effectively processing long sequences of data and avoiding vanishing and exploding gradients. The second BiLSTM layer, consisting of 128 neurons, further mines and abstracts the features output by the first layer, learning more complex feature patterns. This layer captures more advanced fault signatures, improving the model's fault recognition capabilities. The third BiLSTM layer, again consisting of 64 neurons, integrates and refines the features output by the second layer, outputting feature representations suitable for processing by the multi-head attention mechanism.

[0032] The feature representations output by the three-layer BiLSTM are input into a multi-head attention mechanism (four heads). This mechanism enables the model to focus on different aspects of features in different representation subspaces, thereby more comprehensively capturing the correlations between features. Each attention head assigns different attention weights to different parts of the feature by calculating the correlation between the query, key, and value. Specifically, the input feature representation is first projected into the query, key, and value spaces. The similarity score between the query and key is then calculated, converted into attention weights using a softmax function, and the values ​​are weighted and summed according to the attention weights to obtain the output of each attention head. The outputs of the four attention heads are concatenated and linearly transformed to obtain the feature representation processed by the attention mechanism. This feature representation highlights key information relevant to fault diagnosis and suppresses interference from irrelevant information.

[0033] The feature representations processed by the multi-head attention mechanism are input into the output layer of the DNN fault classifier. The output layer assigns corresponding category labels based on the fault type, such as hardware fault, software fault, and network fault. Each broad category further includes specific fault types. The model calculates the degree of match between the input features and each category label and converts this match into a probability value using the softmax function. The category with the highest probability value is the fault type predicted by the model. The predicted probability value is compared with a preset threshold. If the highest probability value exceeds the threshold, the system is deemed to have a corresponding fault. If all probability values ​​are below the threshold, the system is deemed to have no obvious faults.

[0034] The fault classification results are output, clearly indicating whether the system has a fault and the specific fault type. Relevant information about the fault classification is recorded, including the input 128-dimensional feature vector, the model's predicted probability value, and the determined fault type. This information can be used for subsequent analysis and model optimization. Fault information is also synchronized to the system's monitoring and management interface, allowing operations and maintenance personnel to promptly understand system fault conditions and take appropriate measures.

[0035] The training process of the DNN fault classifier model is as follows: Collect a large amount of historical operational data for the Tongxin UOS system, covering both normal system operation and various fault conditions. For example, collect data from the past year related to hardware failures (such as ACPI exceptions and DMA failures), software failures (such as service startup timeouts and dependency loop detection), and network failures (such as DNS pollution and IPsec negotiation failures). Also collect data from normal system operation for comparison. This data includes 28 system metrics collected through system commands and tools, such as the CPU user-mode / kernel-mode time ratio, memory leak detection counters, file system inode error rate, network TCP retransmission rate, etc., as well as system logs, process information, disk information, and network information.

[0036] Label the collected data, clearly indicating the fault type or normal state for each set of data. For example, data from an ACPI failure is labeled "Hardware Fault - ACPI Abnormal"; data from normal operation is labeled "Normal State." This labeling can be performed by professional operations personnel or technical experts to ensure accuracy.

[0037] Remove noise, outliers, and duplicate data from the data. For example, in the system log, there may be some meaningless records due to temporary system fluctuations, or repeated error messages, which need to be cleaned up. Convert the original data into a 128-dimensional feature vector suitable for model input. The specific steps are the same as the previous "Feature Vector Generation" section, including cleaning the log text, tokenizing it, removing stop words, and then calculating the TF-IDF value, constructing the feature vector, and adjusting the dimension to 128 dimensions. For non-text data, such as system indicator data, normalize it and map it to the [0,1] interval to eliminate the dimensional differences between different indicators.

[0038] Divide the preprocessed data into training, validation, and test sets, typically in a ratio of 7:1:2. The training set is used to train the model, allowing it to learn the features and patterns in the data. The validation set is used to evaluate the model's performance during training, adjust its hyperparameters, and prevent overfitting. The test set is used to ultimately evaluate the model's generalization capabilities after training.

[0039] Build a DNN fault classifier model consisting of a 3-layer BiLSTM (64→128→64 neurons) plus a multi-head attention mechanism (4 heads): BiLSTM layer: BiLSTM can capture time series information in data and is suitable for processing data with temporal relationships, such as system operation data. The first BiLSTM layer has 64 neurons, which initially processes the input 128-dimensional feature vector and learns basic feature patterns. The second BiLSTM layer increases to 128 neurons to further mine and abstract features. The third BiLSTM layer returns to 64 neurons to integrate and refine features.

[0040] Multi-head attention mechanism: A multi-head attention mechanism (4 heads) is added after the BiLSTM layer, allowing the model to focus on different aspects of features in different representation subspaces and more comprehensively capture the correlation information between features.

[0041] Output layer: The number of neurons in the output layer is set based on the number of fault types. Each neuron corresponds to a fault type or a normal state. The softmax function is used to convert the output into a probability value, indicating the likelihood that the input data belongs to each fault type.

[0042] The parameters of each model layer (BiLSTM layer, attention layer, and output layer) are assigned random initial values. These parameters are continuously adjusted during training to optimize model performance. A set of data (128-dimensional feature vectors) from the training set is input into the model. The data first passes through three BiLSTM layers. Each BiLSTM layer's neurons perform calculations based on the input and its own weights. A gating mechanism (input gate, forget gate, and output gate) controls the flow and memory of information, and outputs a processed feature representation. The feature representation processed by the BiLSTM layers enters a multi-head attention mechanism. Each attention head projects the input feature representation into the query, key, and value space, calculates the similarity score between the query and key, converts the score into an attention weight using the softmax function, and then performs a weighted sum of the values ​​based on the attention weights to produce the output of each attention head. The outputs of the four attention heads are then concatenated and linearly transformed to produce the feature representation processed by the attention mechanism. Finally, the feature representation processed by the attention mechanism is input into the output layer, which uses the softmax function to calculate the probability of the input data belonging to each fault type.

[0043] The cross-entropy loss function is used to calculate the difference between the probability distribution of the model's output and the annotated true labels. The cross-entropy loss function measures the degree of deviation between the model's predictions and the ground truth. A smaller loss value indicates a more accurate model prediction. Based on the calculated loss value, the backpropagation algorithm is used to calculate the gradient of the loss function with respect to each model parameter. Starting from the output layer, the backpropagation algorithm calculates the gradient layer by layer, propagating the error information back through each layer of the model to determine the contribution of each parameter to the loss. An optimization algorithm (such as the Adam optimization algorithm) is used to update the model parameters based on the calculated gradient. The optimization algorithm adjusts the parameter values ​​based on the magnitude and direction of the gradient, gradually reducing the loss function value and thus improving model performance. This process is repeated for multiple iterations on all the data in the training set until the model performance reaches stability or meets a preset stopping criterion (such as the loss function value no longer decreases or decreases by less than a threshold).

[0044] During model training, regularly evaluate model performance using the validation set. Calculate metrics such as accuracy, recall, and F1 score on the validation set to assess the model's ability to identify different fault types. Based on the validation set evaluation results, adjust the model's hyperparameters, such as the learning rate, batch size, and number of iterations. By continuously trying different hyperparameter combinations, find the optimal hyperparameter settings and improve model performance. After model training is complete and hyperparameters are adjusted, perform a final evaluation of the model using the test set. Calculate various performance metrics on the test set to ensure that the model has good generalization capabilities and can accurately identify fault types in unseen data.

[0045] Obtaining a fault state vector of a Tongxin UOS system and obtaining a fault repair plan according to the fault state vector comprises the following steps: determining a state vector set for representing the fault state of the Tongxin UOS system based on a determined fault type and a fault type-state feature combination mapping table stored in a database; obtaining corresponding features from initial system state indicator data based on the state vector set, and arranging them into a fault state vector of the Tongxin UOS system in a set order; determining a repair plan set corresponding to a current fault state vector based on the determined fault state vector and a fault state vector-repair plan set mapping table stored in a database; evaluating each initial repair plan in the repair plan set using a reinforcement learning reward function to obtain a repair action evaluation value corresponding to each initial repair plan; and defining the initial repair plan corresponding to the maximum repair action evaluation value as the fault repair plan.

[0046] The reinforcement learning reward function is expressed as: ; in, is the output value of the reinforcement learning reward function, is the probability that the initial repair solution a successfully repairs the fault under the current fault state vector s (expressed as the average repair success rate), is the time saved by the initial repair solution a when successfully repairing the fault under the current fault state vector s (the time saved compared to the manual repair time), is the resource usage (the weighted sum of CPU usage and GPU usage) caused by the initial repair solution a successfully repairing the fault under the current fault state vector s, for The weight factor, for The weight factor, for The weight factor of .

[0047] The fault repair plan is verified based on the sandbox environment. If the fault repair plan is feasible, it is executed. If it is not feasible, a new repair plan is generated.

[0048] Traditional system repair tools rely on pre-built rule bases. When faced with new, unknown faults, their repair solutions may not be fully verified, which can be risky. In a sandbox environment, we create an environment that is identical to the UnionTech UOS system and restore it to its pre-fault state. This provides a safe place to verify new repair solutions.

[0049] Create a sandbox environment in the Tongxin UOS system that is completely consistent with the Tongxin UOS system, and restore the system state of the sandbox environment to the same state as before the Tongxin UOS system failure. Follow the operation steps of the fault repair plan and execute repair commands or configuration modifications in the sandbox environment in sequence to obtain multi-dimensional evaluation indicators, including repair success rate, execution time efficiency, and rollback feasibility. Obtain a comprehensive score based on the multi-dimensional evaluation indicators, and verify the feasibility of the fault repair plan based on the comprehensive score.

[0050] Obtain the set threshold stored in the database; if the comprehensive score is greater than or equal to the set threshold and there is no serious system-level risk, the fault repair plan is feasible; if the comprehensive score is less than the set threshold or there is a serious system-level risk, the fault repair plan is not feasible.

[0051] The comprehensive score is obtained as follows: ; in, For the comprehensive score, For the repair success rate, For execution time efficiency, is the actual execution time, is the mean repair time threshold stored in the database, Rollback feasibility is used to evaluate whether the repair operation is reversible, such as whether the modified configuration file has a backup, and whether the updated software version supports downgrade. If the repair fails, it can be quickly rolled back through the backup file or system restore point without affecting the system stability, then R_rollback=1; otherwise, it is assigned a value of 0-1 based on the rollback difficulty, which can be manually evaluated and assigned by experts. for The weight factor, for The weight factor, for The weight factor of .

[0052] By executing the repair plan in a sandbox environment, we can obtain multi-dimensional evaluation indicators such as repair success rate, execution time efficiency, and rollback feasibility. For new and unknown faults, these indicators can more comprehensively evaluate the effectiveness of the repair plan.

[0053] An electronic device includes: a processor; and a memory, wherein computer program instructions are stored in the memory, and when the computer program instructions are executed by the processor, the processor executes the deep learning-based Tongxin UOS system fault rapid repair method as described above.

[0054] A computer-readable storage medium is used to store a program, which, when executed by a processor, implements the deep learning-based Tongxin UOS system fault rapid repair method as described above.

[0055] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0056] The present invention is described with reference to flowcharts and / or block diagrams of systems, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1A device that provides the functions specified in a block or multiple blocks.

[0057] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0058] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0059] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0060] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.

Claims

1. A deep learning-based rapid repair method for Tongxin UOS system faults, characterized by: The following steps are involved: Obtain initial system status indicator data based on system commands and data collection tools; Perform feature engineering on the initial system status indicator data to obtain the Tongxin UOS system feature vector; Determine whether the Tongxin UOS system has a fault and the fault type based on the Tongxin UOS system feature vector, and obtain a fault repair plan after determining that the Tongxin UOS system has a fault; The fault repair plan is verified based on the sandbox environment. If the fault repair plan is feasible, it is executed. If it is not feasible, a new repair plan is generated.

2. The method for rapid repair of Tongxin UOS system faults based on deep learning according to claim 1 is characterized in that: Perform feature engineering on the initial system status indicator data to obtain the Tongxin UOS system feature vector, including the following steps: After cleaning the initial system status indicator data, perform word segmentation using a word segmentation tool to obtain status indicator terms; Calculate the word frequency and inverse document frequency of each state indicator word, and perform weighted summation of the word frequency and inverse document frequency to obtain the TF-IDF value of each state indicator word; Sort the status indicator words in descending order according to the TF-IDF value to obtain the initial Tongxin UOS system feature vector; The dimension of the initial Tongxin UOS system feature vector is adjusted to obtain the Tongxin UOS system feature vector.

3. The method for rapid repair of Tongxin UOS system faults based on deep learning according to claim 2 is characterized in that: The dimension of the initial UOS system feature vector is adjusted to obtain the UOS system feature vector, including the following steps: Determine whether the eigenvector dimension of the initial Tongxin UOS system eigenvector is greater than the set K value: If the dimension of the feature vector is greater than the set K value, principal component analysis is performed on each feature in the initial UOS system feature vector, and the state indicator words corresponding to each feature are reordered according to the variance contribution, and the top K state indicator words are obtained to construct the UOS system feature vector; If the dimension of the feature vector is not greater than the set K value, zero value padding is performed to obtain the UOS system feature vector with a feature vector dimension of K.

4. The method for rapid repair of Tongxin UOS system faults based on deep learning according to claim 1 is characterized in that: Obtain a troubleshooting solution, including the following steps: Input the Tongxin UOS system feature vector into the trained fault type diagnosis model to determine whether there is a fault in the Tongxin UOS system and the fault type; If there is a fault in the Tongxin UOS system, the fault state vector of the Tongxin UOS system is obtained according to the determined fault type; The fault repair plan is obtained based on the fault state vector of the Tongxin UOS system.

5. The method for rapid repair of Tongxin UOS system faults based on deep learning according to claim 4 is characterized in that: The fault type diagnosis model is a BiLSTM-multi-head attention DNN fault classifier model, where the BiLSTM has 3 layers and the multi-head attention has 4 heads: The first layer of BiLSTM includes 64 neurons, the second layer of BiLSTM includes 128 neurons, and the third layer of BiLSTM includes 64 neurons.

6. The method for rapid repair of Tongxin UOS system faults based on deep learning according to claim 4 is characterized in that: Obtain the fault state vector of the Tongxin UOS system and obtain the fault repair plan based on the fault state vector, including the following steps: Determine a state vector set for representing the fault state of the Tongxin UOS system based on the determined fault type and the fault type-state feature combination mapping table stored in the database; Based on the state vector set, the corresponding features are obtained from the initial system state indicator data and arranged in the set order to form the fault state vector of the Tongxin UOS system; Determine a repair solution set corresponding to the current fault state vector based on the determined fault state vector and a fault state vector-repair solution set mapping table stored in a database; The reinforcement learning reward function is used to evaluate each initial repair plan in the repair plan set to obtain the repair action evaluation value corresponding to each initial repair plan; The initial repair plan corresponding to the maximum repair action evaluation value is defined as the fault repair plan.

7. The method for rapid repair of Tongxin UOS system faults based on deep learning according to claim 6 is characterized in that: The reinforcement learning reward function is expressed as: ; in, is the output value of the reinforcement learning reward function, is the probability that the initial repair solution a successfully repairs the fault under the current fault state vector s, is the time saved by the initial repair solution a to successfully repair the fault under the current fault state vector s, is the resource occupancy rate caused by the initial repair solution a successfully repairing the fault under the current fault state vector s, for The weight factor, for The weight factor, for The weight factor of .

8. The method for rapid repair of Tongxin UOS system faults based on deep learning according to claim 1 is characterized in that: Verify the fault repair solution in a sandbox environment, including the following steps: Create a sandbox environment in the Tongxin UOS system that is completely consistent with the Tongxin UOS system, and restore the system state of the sandbox environment to the same state as before the Tongxin UOS system failure; Follow the troubleshooting plan's steps and execute repair commands or configuration changes in a sandbox environment. Obtain multi-dimensional evaluation metrics, including repair success rate, execution time efficiency, and rollback feasibility. A comprehensive score is obtained based on multi-dimensional evaluation indicators, and the feasibility of the fault repair plan is verified based on the comprehensive score.

9. The method for rapid repair of Tongxin UOS system faults based on deep learning according to claim 8 is characterized in that: The feasibility of the fault repair solution is verified based on the comprehensive score, including the following steps: Get the set threshold stored in the database; If the comprehensive score is greater than or equal to the set threshold and there is no serious system-level risk, the fault repair plan is feasible; If the comprehensive score is less than the set threshold or there is a serious system-level risk, the fault repair plan is not feasible.

10. The method for rapid repair of Tongxin UOS system faults based on deep learning according to claim 8 is characterized in that: The method for obtaining the comprehensive score is as follows: ; in, For the comprehensive score, For the repair success rate, For execution time efficiency, is the actual execution time, is the mean repair time threshold stored in the database, For rollback feasibility, for The weight factor, for The weight factor, for The weight factor of .

Citation Information

Patent Citations

  • System integration fault diagnosis system

    CN117076175A

  • Distributed service self-healing method and device

    CN118193267A

  • Fault analysis method for automatic testing equipment of vehicle machine

    CN119537079A

  • Storage hard disk remote diagnosis system and method based on Internet of Things

    CN120256178A