A fault prediction and repair method, system, electronic device and storage medium
By using embedded KVM and AI visual recognition technology, the system automatically identifies the server's operating status and enables remote control, solving the problem of low efficiency caused by traditional server maintenance relying on manual operation, and achieving efficient fault prediction and repair.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 广东美凯技术有限公司
- Filing Date
- 2025-07-08
- Publication Date
- 2026-05-12
AI Technical Summary
Traditional server maintenance relies on manual operation, which is inefficient and cannot effectively identify and respond to server crashes or BIOS malfunctions.
The system uses embedded KVM to acquire image streams from the server terminal, uses AI vision to identify the operating status, constructs a status graph, and performs remote control operations based on this mapping processing strategy, while also providing feedback verification and optimization adjustments.
It enables automated identification and efficient repair of server faults, improving the server's ability to automatically identify, respond to, and repair faults.
Smart Images

Figure CN121008858B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a fault prediction and repair method, system, electronic device and storage medium. Background Technology
[0002] In traditional server maintenance, administrators rely on operating system logs, SNMP alerts, or manual observation of remote desktops to identify faults. However, these traditional methods fail in scenarios such as operating system crashes, system freezes, or BIOS malfunctions. While some communication technologies enable remote access to servers, they lack intelligent judgment and automated response mechanisms, still requiring manual analysis and judgment. Therefore, current technologies are inefficient for server maintenance. Summary of the Invention
[0003] The main objective of this application is to provide a fault prediction and repair method, system, electronic device, and storage medium, which aims to solve at least one problem of the prior art.
[0004] To achieve the above objectives, one aspect of this application proposes a fault prediction and repair method, the method comprising:
[0005] The image stream of the server terminal is acquired using an embedded KVM; the image stream includes multiple acquired images.
[0006] AI visual recognition is performed on the acquired images to obtain the operating status of the server terminal, and a status graph information is constructed based on the operating status.
[0007] The processing strategy is determined based on the mapping of state diagram information.
[0008] Based on the processing strategy, embedded KVM is used to remotely control the server terminal.
[0009] Feedback verification and optimization adjustments to the AI visual recognition are based on the results of remote control operations.
[0010] In some embodiments, acquiring the image stream of a server terminal using an embedded KVM includes the following steps:
[0011] Using embedded KVM, the raw image stream of the server terminal is acquired based on a preset frequency through a high-speed frame buffering mechanism;
[0012] Each original acquired image in the original image stream is preprocessed and then organized to obtain the target image stream;
[0013] Image preprocessing includes color space conversion and image enhancement.
[0014] In some embodiments, the operating state includes a stage state and a timing state. The operating state of the server terminal is obtained by performing AI visual recognition based on the acquired images, including the following steps:
[0015] Based on the acquired images, a multi-task convolutional neural network is used to perform state classification and recognition to obtain the stage state corresponding to the acquired images.
[0016] The multi-task convolutional neural network includes a lightweight backbone, a stage classification branch, and an error type recognition branch. It is used to perform state classification and recognition to obtain the stage state corresponding to the acquired image, including the following steps:
[0017] A lightweight backbone is used to extract features from the acquired images to obtain multi-level image features;
[0018] Based on multi-level image features, stage classification is output using stage classification branch mapping.
[0019] When a stage is classified as an error stage, the error type is output by using error type identification branch mapping based on multi-level image features.
[0020] Key text recognition is performed on the acquired images to obtain text information;
[0021] The abnormal state of the server terminal is determined by looking up the table based on the error type and text information, and this is used as the stage state.
[0022] A time-series convolutional network is used to determine the temporal state of the image sequence of consecutive frames in an image stream.
[0023] In some embodiments, constructing state graph information based on runtime status includes the following steps:
[0024] State data of the server terminal is collected based on a state diagram model using a finite state machine.
[0025] The status diagram information is obtained by integrating the running status and status data.
[0026] In some embodiments, determining a processing strategy based on state diagram information mapping includes the following steps:
[0027] Based on state diagram information, a processing strategy is output through neural network mapping that integrates rules;
[0028] Among them, the neural network that integrates rules has pre-learned the mapping relationship between different state diagram information and preset business rules.
[0029] In some embodiments, based on a processing strategy, remote control operations on the server terminal are performed using an embedded KVM, including the following steps:
[0030] Based on the processing strategy, the virtual USB technology of embedded KVM is used to call the embedded low-level power control API to trigger target control commands to remotely control the server terminal.
[0031] In some embodiments, feedback verification and optimization adjustments to AI visual recognition based on the results of remote control operations include the following steps:
[0032] In response to the completion signal of the remote control operation, the feedback verification mechanism is initiated to re-acquire the image of the server terminal using embedded KVM and perform AI visual recognition to obtain the processing result of the server terminal.
[0033] The operation result of the remote control operation is determined based on the processing result. When the operation result is operation failure, the control server terminal starts the rollback mechanism or adjusts the recognition confidence threshold of the model called by the AI visual recognition.
[0034] The recognition results corresponding to AI visual recognition and the operation logs and processing results corresponding to remote control operations are uploaded to the training dataset on the cloud server.
[0035] The training dataset is used to iteratively optimize the model invoked by the AI visual recognition system.
[0036] To achieve the above objectives, another aspect of this application proposes a fault prediction and repair system, the system comprising:
[0037] The data acquisition module is used to acquire image streams from the server terminal using an embedded KVM; the image stream includes multiple acquired images.
[0038] The visual recognition module is used to perform AI visual recognition based on the acquired images, obtain the operating status of the server terminal, and construct a status diagram based on the operating status.
[0039] The strategy mapping module is used to determine the processing strategy based on the state diagram information mapping;
[0040] The remote control module is used to remotely control the server terminal using an embedded KVM based on the processing strategy.
[0041] The feedback optimization module is used to perform feedback verification and optimize the AI visual recognition based on the results of remote control operations.
[0042] To achieve the above objectives, another aspect of the embodiments of this application proposes an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the aforementioned method.
[0043] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned method.
[0044] To achieve the above objectives, another aspect of this application provides a computer program product, including a computer program that, when executed by a processor, implements the aforementioned method.
[0045] The embodiments of this application include at least the following beneficial effects: This application provides a fault prediction and repair method, system, electronic device, storage medium, and program product. This solution utilizes an embedded KVM to acquire an image stream from a server terminal; wherein the image stream includes multiple acquired images; AI visual recognition is performed based on the acquired images to obtain the operating status of the server terminal; a state graph is constructed based on the operating status; a processing strategy is determined based on the state graph information mapping; based on the processing strategy, the embedded KVM is used to remotely control the server terminal; feedback verification and optimization adjustments are made based on the results of the remote control operations. This application, based on image-level remote access achieved through embedded KVM, further introduces AI visual recognition and strategy mapping to achieve remote control operations. This application can effectively improve the automatic identification, response, and repair capabilities of servers in fault states. This application can automatically and efficiently achieve server fault prediction and repair. Attached Figure Description
[0046] Figure 1 This is a schematic diagram of an implementation environment for the fault prediction and repair method provided in this application embodiment;
[0047] Figure 2 This is a flowchart illustrating a fault prediction and repair method provided in an embodiment of this application;
[0048] Figure 3 This is a schematic diagram illustrating the architecture and principle of the fault prediction and repair method provided in the embodiments of this application;
[0049] Figure 4 This is a schematic diagram of the strategy decision-making process of the fault prediction and repair method provided in the embodiments of this application;
[0050] Figure 5 This is a schematic diagram of the structure of a fault prediction and repair system provided in an embodiment of this application;
[0051] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0052] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of systems and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.
[0053] It is understood that the terms “first,” “second,” etc., used in this application may be used herein to describe various concepts, but unless otherwise stated, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the words “if,” “when,” or “in response to a determination” as used herein may be interpreted as “when…” or “when…” or “in response to a determination.”
[0054] As used in this application, the terms "at least one", "multiple", "each", "any", etc., "at least one" includes one, two or more, "multiple" includes two or more, "each" refers to each of the corresponding multiples, and "any" refers to any one of the multiples.
[0055] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0056] In related technologies, during traditional server maintenance, administrators rely on operating system layer logs, SNMP protocol alarms, or manual observation of remote desktops to identify faults and perform manual repairs, which is inefficient.
[0057] In view of this, this application provides a fault prediction and repair method. This method utilizes an embedded KVM to acquire an image stream from a server terminal; the image stream includes multiple acquired images; AI visual recognition is performed based on the acquired images to obtain the server terminal's operating status; a status graph is constructed based on the operating status; a processing strategy is determined based on the status graph information; based on the processing strategy, the embedded KVM is used to remotely control the server terminal; feedback verification and optimization adjustments are made based on the results of the remote control operations. This application, based on image-level remote access achieved through embedded KVM, further introduces AI visual recognition and strategy mapping to achieve remote control operations. This application can effectively improve the server's automatic identification, response, and repair capabilities under fault conditions. This application can automatically and efficiently achieve server fault prediction and repair.
[0058] It is understood that the fault prediction and repair method provided in this application can be applied to any computer device with data processing and computing capabilities, and this computer device can be various terminals or servers. When the computer device in the embodiment is a server, the server is an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. Optionally, the terminal can be a smartphone, tablet, laptop, or desktop computer, but it is not limited to these.
[0059] like Figure 1 The diagram shown is a schematic representation of an implementation environment provided in an embodiment of this application. (Refer to...) Figure 1 The implementation environment includes at least one terminal 102 and a server 101. The terminal 102 and the server 101 can be connected via a network, either wirelessly or via a wired connection, to complete data transmission and exchange.
[0060] Server 101 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0061] Additionally, server 101 can also be a node server in a blockchain network. Blockchain is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms.
[0062] Terminal 102 can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, etc., but is not limited to these. Terminal 102 and server 101 can be directly or indirectly connected via wired or wireless communication, and this embodiment of the application does not impose any limitations.
[0063] For example, based on Figure 1 The implementation environment shown in this application embodiment provides a fault prediction and repair method. The following description uses the application of the fault prediction and repair method in server 101 as an example. It can be understood that the fault prediction and repair method can also be applied to terminal 102. Specifically, terminal 102 or server 101 can be used to execute the relevant data processing logic of the method of this application to implement the corresponding method flow.
[0064] Reference Figure 2 , Figure 2 This is an optional flowchart of the fault prediction and repair method provided in the embodiments of this application. The subject executing the fault prediction and repair method can be any of the aforementioned computer devices (including servers or terminals). Figure 2 The method may include, but is not limited to, steps S100 to S500.
[0065] Step S100: Use embedded KVM to obtain the image stream from the server terminal;
[0066] The image stream includes multiple captured images;
[0067] It should be noted that in some embodiments, step S100 may include the following steps: using an embedded KVM, acquiring the original image stream of the server terminal based on a preset frequency through a high-speed frame buffer mechanism; performing image preprocessing on each original acquired image in the original image stream to obtain the target image stream; wherein, image preprocessing includes color space conversion and image enhancement processing.
[0068] For example, in some specific implementations, an embedded KVM module is used to acquire image streams from the server terminal at a frequency of 2–5 frames per second through a high-speed frame buffering mechanism. The KVM module is compatible with VGA and HDMI input interfaces and supports a maximum resolution of 3840x2160. The acquired raw images undergo color space conversion (unifying the BT.709 color gamut) and image enhancement processing (including noise reduction, sharpening, and brightness and contrast adjustment), effectively improving the accuracy of subsequent visual recognition.
[0069] Step S200: Perform AI visual recognition based on the acquired images to obtain the operating status of the server terminal, and construct a status diagram based on the operating status;
[0070] It should be noted that the operating status includes stage status and timing status. In some embodiments, the operating status of the server terminal is obtained by performing AI visual recognition based on the acquired images, which may include the following steps:
[0071] Regarding the stage status:
[0072] Based on the acquired images, a multi-task convolutional neural network is used to perform state classification and recognition to obtain the stage state corresponding to the acquired images.
[0073] The multi-task convolutional neural network includes a lightweight backbone, a stage classification branch, and an error type recognition branch. It is used to perform state classification and recognition to obtain the stage state corresponding to the acquired image, including the following steps:
[0074] A lightweight backbone is used to extract features from the acquired images to obtain multi-level image features;
[0075] Based on multi-level image features, stage classification is output using stage classification branch mapping.
[0076] When a stage is classified as an error stage, the error type is output by using error type identification branch mapping based on multi-level image features.
[0077] Key text recognition is performed on the acquired images to obtain text information;
[0078] The abnormal state of the server terminal is determined by looking up the table based on the error type and text information, and this is used as the stage state.
[0079] For example, in some specific implementations, a multi-task convolutional neural network (Multi-taskCNN) structure can be used to perform high-precision stage classification on the acquired images, accurately distinguishing whether the server is in a state such as POST, self-test, BIOS settings, operating system loading, blue screen, system crash, or hard drive error. The network structure uses a lightweight backbone (such as MobileNetV3) to extract multi-level image features and sets multiple task output branches, including a stage classification branch, an error type identification branch, and a confidence score branch (used to determine the confidence score of the stage classification branch for each stage classification and the confidence score of the error type identification branch for each error type; usually, the stage classification and error type with the highest confidence are taken as the result of the corresponding branch), realizing parallel inference and collaborative state judgment within the same network. Taking the "blue screen crash" scenario as an example, the system acquired images with typical large areas of blue background and key text features (such as "Your PC ran into an aproblem"). After the image is input into the neural network, the stage classification branch identifies it as an "error stage," and the error type branch further identifies it as a "blue screen." Combined with semantic analysis of the text region in the image, the system ultimately achieves an accurate determination of the server's abnormal state. Specifically, key text recognition in semantic analysis applications can employ an OCR model based on the Transformer architecture to efficiently identify key prompt text information on the screen, including system error messages, BIOS interactive prompts, warnings, and status feedback (such as "No boot device" and "Press F1 to continue"). The model has strong contextual understanding capabilities and high robustness to blurred, partially occluded, or complex background text.
[0080] In some specific application scenarios, the stage classification branch can use global feature compression (global average pooling) to adapt to the whole image state recognition; the error recognition branch can use spatial attention mechanism to focus on local abnormal regions (such as the blue screen text area); and the confidence branch can use dual-path heterogeneous convolution to adapt to the feature scale of different tasks.
[0081] Regarding timing states:
[0082] A time-series convolutional network is used to determine the temporal state of consecutive frames of captured images in an image stream.
[0083] For example, in some specific implementations, temporal state judgment can be achieved by introducing a Temporal Series Convolutional Network (TSCN) to model and analyze continuous frame image sequences. This can be used to identify whether the server is in an abnormal loop or has been unresponsive for a long time, such as repeated restarts, the system being stuck in the BIOS, or the loading screen exceeding a threshold time. The TSCN network extracts the temporal features of image frames (such as color distribution, UI component position changes, text dynamics, etc.) and combines them with a sliding window mechanism to determine the state evolution trend. Taking the identification of "repeated restarts" anomalies as an example, the continuously acquired image sequence of the system exhibits periodic screen switching characteristics (such as POST self-test screen -> LOGO startup screen -> black screen -> POST again). After inputting the image sequence within the time window into the TSCN network, the model detects that the screen changes show a periodic oscillation trend and that a stable interface (such as the operating system desktop or login screen) has not been entered within the set time window. Based on this, a "high repetition periodic feature" judgment is made, triggering a policy response.
[0084] It should also be noted that all the visual AI models mentioned in the aforementioned AI visual recognition applications are uniformly converted into a standard network format and deployed to the NPU chip built into KVM for inference computation, realizing low-latency (millisecond-level) real-time inference processing at the edge.
[0085] In some embodiments, constructing state diagram information based on the running state may include the following steps: collecting state data of the server terminal based on a state diagram model of a finite state machine; and integrating the running state and state data to obtain state diagram information.
[0086] For example, in some specific implementations, a state diagram model based on a finite state machine (FSM) can be constructed to track changes in the server's operating state in real time, and the server system state diagram can be updated and maintained in real time by combining AI visual recognition results.
[0087] In some specific application scenarios, the collected state data and the operational status of the AI visual recognition results can be used as input to the FSM (Finite State Machine) model. The FSM model then uses predefined state transition rules to determine the current state of the server system (e.g., normal operation, high load, potential fault, hardware failure, network interruption, maintenance, etc.). Based on the FSM's determination, a state diagram reflecting the current and historical state evolution of the entire server system is updated and maintained in real time.
[0088] Step S300: Determine the processing strategy based on the state diagram information mapping;
[0089] It should be noted that in some embodiments, step S300 may include the following steps: based on state graph information, outputting a processing strategy through a neural network mapping of fusion rules; wherein, the neural network of fusion rules has pre-learned the mapping relationship between different state graph information and preset business rules.
[0090] For example, in some specific implementations, a Hybrid Rule-MLP can be used to accurately output the optimal processing strategy for different fault categories (such as type A: system freeze, type B: blue screen crash, type C: waiting for user input) by fusing automatically generated state diagram information with preset business rules, and automatically generate a tree-like processing path for the fault.
[0091] Step S400: Based on the processing strategy, remote control operation is performed on the server terminal using embedded KVM;
[0092] It should be noted that in some embodiments, step S400 may include the following steps: based on the processing strategy, using the virtual USB technology of embedded KVM to call the embedded underlying power control API to trigger the target control command to perform remote control operation on the server terminal.
[0093] For example, in some specific implementations, virtual USB interface technology based on embedded KVM is used to simulate precise remote keyboard and mouse operations, supporting precise commands such as F1, DEL, arrow keys, Enter key, and ESC key to complete automatic BIOS settings and system recovery. Simultaneously, by calling the embedded low-level power control API, control commands such as Reset, PowerCycle, and ForceShutdown are sent to achieve precise and rapid remote control operations at the server hardware level.
[0094] Step S500: Based on the results of the remote control operation, feedback verification is performed and the AI visual recognition is optimized and adjusted.
[0095] It should be noted that, in some embodiments, step S500 may include the following steps: in response to the completion signal of the remote control operation, a feedback verification mechanism is initiated to re-acquire the image of the server terminal using embedded KVM and perform AI visual recognition to obtain the processing result of the server terminal; based on the processing result, the operation result of the remote control operation is determined; if the operation result is an operation failure, the server terminal is controlled to initiate a rollback mechanism or adjust the recognition confidence threshold of the model called by the AI visual recognition; the recognition result corresponding to the AI visual recognition and the operation log and processing result corresponding to the remote control operation are uploaded to the training dataset of the cloud server; wherein, the training dataset is used to iteratively optimize the model called by the AI visual recognition.
[0096] For example, in some specific implementations, after the fault handling operation is completed, a feedback verification mechanism is immediately initiated to re-acquire the server terminal image and use the AI visual recognition model to confirm the processing result. If the recognition determines that the operation has failed, the system automatically initiates a rollback mechanism or adjusts the AI model's recognition confidence threshold to avoid repeated misjudgments. At the same time, the system automatically records complete processing data, including operation logs, recognition results, and processing effects, and periodically feeds it back to the cloud server to expand the training dataset, enabling continuous adaptive optimization and iterative improvement of the AI model.
[0097] To explain in detail the principles of the technical solution of this application, the overall process of this application will be described below with reference to some specific embodiments. It is easy to understand that the following is an explanation of the technical principles of this application and should not be regarded as a limitation of this application.
[0098] In view of the relevant shortcomings of the prior art, embodiments of this application provide a fault prediction and repair method, such as Figure 3 and Figure 4 As shown, the method of this application can be implemented through the following process:
[0099] (I) Image Acquisition and Preprocessing:
[0100] An embedded KVM module is used to acquire image streams from the server terminal at a frequency of 2–5 frames per second via a high-speed frame buffering mechanism. The KVM module is compatible with VGA and HDMI input interfaces and supports a maximum resolution of 3840x2160. The acquired raw images undergo color space conversion (unifying the BT.709 color gamut) and image enhancement processing (including noise reduction, sharpening, and brightness and contrast adjustment) to effectively improve the accuracy of subsequent visual recognition.
[0101] (II) AI Visual State Recognition:
[0102] State classification and recognition:
[0103] A multi-task convolutional neural network (CNN) structure is employed to perform high-precision stage classification on acquired images, accurately distinguishing between server states such as POST (Power-On Self-Test), BIOS setup, operating system loading, blue screen, system crash, or hard drive error. The network structure uses a lightweight backbone (such as MobileNetV3) to extract multi-level image features and sets multiple task output branches, including a stage classification branch, an error type identification branch, and a confidence score branch (used to determine the confidence scores of the stage classification branch for each stage category and the error type identification branch for each error type; typically, the stage category and error type with the highest confidence score are taken as the result of the corresponding branch). This enables parallel inference and collaborative state judgment within the same network. Taking the "blue screen crash" scenario as an example, the system-acquired image typically has a large blue background and key text features (such as "Your PC ran into a problem"). After the image is input into the neural network, the stage classification branch identifies it as an "error stage," and the error type branch further identifies it as a "blue screen." Combined with semantic analysis of the text regions in the image, the system ultimately achieves an accurate determination of the server's abnormal state.
[0104] Key text recognition:
[0105] The OCR model, employing the Transformer architecture, efficiently recognizes key on-screen prompts, including system error messages, BIOS interactive prompts, warnings, and status feedback (such as "No boot device" and "Press F1 to continue"). The model possesses strong contextual understanding capabilities and is highly robust to blurry, partially occluded, or complex background text.
[0106] Timing state determination:
[0107] A Temporal Series Convolutional Network (TSCN) is introduced to model and analyze continuous frame image sequences to identify whether a server is in an abnormal loop or unresponsive state for a long time, such as repeated restarts, system stuck in BIOS, or loading screens exceeding a threshold time. The TSCN network extracts temporal features of image frames (such as color distribution, UI component position changes, and text dynamics) and combines them with a sliding window mechanism to determine the state evolution trend. Taking the identification of "repeated restarts" anomalies as an example, the continuously acquired image sequence of the system exhibits periodic screen switching characteristics (e.g., POST self-test screen -> LOGO startup screen -> black screen -> POST again). After inputting the image sequence within the time window into the TSCN network, the model detects that the screen changes show a periodic oscillation trend and that a stable interface (such as the operating system desktop or login screen) has not been entered within the set time window. Therefore, it determines and makes a "high repetition periodic feature" judgment, triggering a policy response.
[0108] All visual AI models are uniformly converted into a standard network format and deployed to the NPU chip built into KVM for inference computation, realizing low-latency (millisecond-level) real-time inference processing at the edge.
[0109] (III) Status Assessment and Strategy Decision-Making:
[0110] State machine modeling:
[0111] A state diagram model based on a finite state machine (FSM) is constructed to track changes in the server's operating state in real time. Combined with AI visual recognition results, the server system state diagram is updated and maintained in real time.
[0112] Intelligent strategy decision-making:
[0113] The Hybrid Rule-MLP is designed to integrate automatically generated state graph information with preset business rules, accurately output the optimal handling strategy for different fault categories (such as A: system freeze, B: blue screen crash, C: waiting for user input), and automatically generate a tree-like processing path for the fault.
[0114] (iv) Fault Handling and Remote Control Commands:
[0115] Based on embedded KVM virtual USB interface technology, it simulates precise remote keyboard and mouse operation, supports precise commands such as F1, DEL, arrow keys, Enter key, ESC key, etc., and completes automatic BIOS settings and system recovery.
[0116] Meanwhile, by calling the embedded low-level power control API, control commands such as Reset, PowerCycle, and ForceShutdown can be sent to accurately and quickly realize remote control operations at the server hardware level.
[0117] (V) Feedback Verification and Continuous Learning:
[0118] After the fault handling operation is completed, the feedback verification mechanism is immediately activated to re-acquire the server terminal image and use the AI visual recognition model to confirm the processing result.
[0119] If the identification determines the operation has failed, the system automatically initiates a rollback mechanism or adjusts the AI model's recognition confidence threshold to avoid repeated misjudgments. Simultaneously, the system automatically records complete processing data, including operation logs, recognition results, and processing effects, and periodically feeds this data back to the cloud server to expand the training dataset, enabling continuous adaptive optimization and iterative improvement of the AI model.
[0120] By implementing the above technical solutions, this system effectively solves the problems of passive maintenance and high operation and maintenance costs in server management, and comprehensively improves the reliability, maintainability and intelligence level of the server.
[0121] In summary, the method and technology approach of this application focuses on a closed-loop process of "perception-judgment-execution-feedback," with the core employing embedded KVM image stream coupled with multimodal AI algorithms to achieve visualized intelligent control across operating system layers. The purpose of this application is to address the problems of traditional remote server maintenance, such as heavy reliance on manual intervention, inability to identify system crashes or BIOS stagnation states, and slow response times. By introducing AI visual recognition and intelligent policy control mechanisms, it enhances the server's automatic identification, response, and repair capabilities under fault conditions.
[0122] like Figure 5 As shown in the figure, this application embodiment also provides a fault prediction and repair system 900, which can implement the above-mentioned method. The system includes:
[0123] The data acquisition module 901 is used to acquire the image stream of the server terminal using an embedded KVM; wherein the image stream includes multiple acquired images;
[0124] The visual recognition module 902 is used to perform AI visual recognition based on the acquired images, obtain the operating status of the server terminal, and construct a status diagram information based on the operating status.
[0125] Strategy mapping module 903 is used to determine the processing strategy based on the state diagram information mapping;
[0126] The remote control module 904 is used to remotely control the server terminal using an embedded KVM based on a processing strategy.
[0127] The feedback optimization module 905 is used to perform feedback verification and optimize and adjust the AI visual recognition based on the results of remote control operations.
[0128] It is understood that the content of the above method embodiments is applicable to this system embodiment. The specific functions implemented in this system embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0129] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0130] It is understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0131] like Figure 6 As shown, Figure 6 The hardware structure of an electronic device 1000 according to another embodiment is illustrated. The electronic device 1000 includes:
[0132] The processor 1001 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (aSIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.
[0133] The memory 1002 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RaM). The memory 1002 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1002 and is called and executed by the processor 1001 using the network node population optimization method of the embodiments of this application.
[0134] Input / output interface 1003 is used to implement information input and output;
[0135] The communication interface 1004 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0136] Bus 1005 transmits information between various components of the device (e.g., processor 1001, memory 1002, input / output interface 1003, and communication interface 1004);
[0137] The processor 1001, memory 1002, input / output interface 1003 and communication interface 1004 are connected to each other within the device via bus 1005.
[0138] The electronic device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0139] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.
[0140] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0141] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0142] It is understood that the content of the above method embodiments is applicable to the embodiments of this program product. The specific functions implemented by the embodiments of this program product are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0143] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0144] The fault prediction and repair method, system, electronic device, storage medium, and program product provided in this application acquire an image stream from a server terminal using an embedded KVM. The image stream includes multiple acquired images. AI visual recognition is performed based on the acquired images to obtain the server terminal's operating status, and a status graph is constructed based on the operating status. A processing strategy is determined based on the status graph information. Based on the processing strategy, the embedded KVM is used to remotely control the server terminal. Feedback verification and optimization adjustments are made based on the results of the remote control operations. This application, building upon image-level remote access achieved through embedded KVM, further introduces AI visual recognition and strategy mapping to realize remote control operations. This application effectively improves the server's automatic identification, response, and repair capabilities under fault conditions. This application can automatically and efficiently achieve server fault prediction and repair.
[0145] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0146] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0147] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0148] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0149] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0150] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0151] In the embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between systems or units may be electrical, mechanical, or other forms.
[0152] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0153] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0154] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0155] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A fault prediction and repair method, characterized in that, The method includes the following steps: An image stream from a server terminal is acquired using an embedded KVM; wherein the image stream includes multiple acquired images. AI visual recognition is performed on the acquired images to obtain the operating status of the server terminal, and a status graph information is constructed based on the operating status. The processing strategy is determined based on the state diagram information mapping. Based on the processing strategy, the embedded KVM is used to remotely control the server terminal. Feedback verification and optimization adjustments are performed based on the results of the remote control operation; The operating status includes stage status and time sequence status. The step of performing AI visual recognition based on the acquired images to obtain the operating status of the server terminal includes the following steps: Based on the acquired images, a multi-task convolutional neural network is used to perform state classification and recognition to obtain the stage state corresponding to the acquired images. The multi-task convolutional neural network includes a lightweight backbone, a stage classification branch, and an error type recognition branch. The step of using the multi-task convolutional neural network to perform state classification and recognition to obtain the stage state corresponding to the acquired image includes the following steps: The lightweight backbone is used to extract features from the acquired images to obtain multi-level image features; wherein, based on the confidence scores of each stage classification branch for each stage classification and the confidence scores of each error type recognition branch for each error type, the stage classification and error type with the highest confidence are taken as the results of the corresponding branches. Based on the multi-level image features, the stage classification branch mapping is used to output the stage classification; When the stage is classified as an error stage, the error type is output by using the error type identification branch mapping based on the multi-level image features. Key text recognition is performed on the acquired images to obtain text information; The abnormal state of the server terminal is determined by looking up the table based on the error type and the text information, and this abnormal state is taken as the stage state. The temporal state is determined by using a time-series convolutional network to analyze the image sequence of consecutive frames in the image stream.
2. The method according to claim 1, characterized in that, The method of acquiring the image stream from the server terminal using embedded KVM includes the following steps: Using the embedded KVM, the original image stream of the server terminal is acquired at a preset frequency through a high-speed frame buffering mechanism; Each original acquired image in the original image stream is preprocessed to obtain the target image stream; The image preprocessing includes color space conversion and image enhancement processing.
3. The method according to claim 1, characterized in that, The process of constructing the state diagram information based on the operating state includes the following steps: The state data of the server terminal is collected based on the state diagram model of a finite state machine. The state diagram information is obtained by integrating the operating status and the status data.
4. The method according to claim 1, characterized in that, The process of determining the processing strategy based on the state diagram information mapping includes the following steps: Based on the state diagram information, the processing strategy is output through a neural network mapping that integrates the rules; The neural network for the fusion rules has pre-learned the mapping relationship between different state diagram information and preset business rules.
5. The method according to claim 1, characterized in that, The method of remotely controlling the server terminal using the embedded KVM based on the processing strategy includes the following steps: Based on the aforementioned processing strategy, the embedded KVM's virtual USB technology is used to call the embedded underlying power control API to trigger target control commands to perform the remote control operation on the server terminal.
6. The method according to any one of claims 1 to 5, characterized in that, The feedback verification and optimization adjustment of the AI visual recognition based on the results of the remote control operation includes the following steps: In response to the completion signal of the remote control operation, a feedback verification mechanism is initiated to re-acquire the image of the server terminal using embedded KVM and perform the AI visual recognition to obtain the processing result of the server terminal. Based on the processing result, the operation result of the remote control operation is determined. When the operation result is operation failure, the server terminal is controlled to start the rollback mechanism or the recognition confidence threshold of the model called by the AI visual recognition is adjusted. The recognition results corresponding to the AI visual recognition, the operation logs corresponding to the remote control operation, and the processing results are uploaded to the training dataset of the cloud server. The training dataset is used to iteratively optimize the model invoked by the AI visual recognition.
7. A fault prediction and repair system, characterized in that, The system includes: A data acquisition module is used to acquire image streams from a server terminal using an embedded KVM; wherein the image stream includes multiple acquired images; The visual recognition module is used to perform AI visual recognition based on the acquired images to obtain the operating status of the server terminal, and to construct a status graph information based on the operating status. The strategy mapping module is used to determine the processing strategy based on the state diagram information mapping; A remote control module is used to remotely control the server terminal using the embedded KVM based on the processing strategy. The feedback optimization module is used to perform feedback verification and optimize the AI visual recognition based on the results of the remote control operation. The operating status includes stage status and time sequence status. The step of performing AI visual recognition based on the acquired images to obtain the operating status of the server terminal includes the following steps: Based on the acquired images, a multi-task convolutional neural network is used to perform state classification and recognition to obtain the stage state corresponding to the acquired images. The multi-task convolutional neural network includes a lightweight backbone, a stage classification branch, and an error type recognition branch. The step of using the multi-task convolutional neural network to perform state classification and recognition to obtain the stage state corresponding to the acquired image includes the following steps: The lightweight backbone is used to extract features from the acquired images to obtain multi-level image features; wherein, based on the confidence scores of each stage classification branch for each stage classification and the confidence scores of each error type recognition branch for each error type, the stage classification and error type with the highest confidence are taken as the results of the corresponding branches. Based on the multi-level image features, the stage classification branch mapping is used to output the stage classification; When the stage is classified as an error stage, the error type is output by using the error type identification branch mapping based on the multi-level image features. Key text recognition is performed on the acquired images to obtain text information; The abnormal state of the server terminal is determined by looking up the table based on the error type and the text information, and this abnormal state is taken as the stage state. The temporal state is determined by using a time-series convolutional network to analyze the image sequence of consecutive frames in the image stream.
8. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method according to any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 6.