Automatic inspection method and system for server

Through the integrated server automatic inspection methods and systems of automated task orchestration, dynamic grouping management and intelligent inspection, the problems of inefficiency and poor flexibility in the traditional server management mode are solved, and efficient management and intelligent inspection of large-scale servers are achieved.

CN120144408APending Publication Date: 2025-06-13CLP CHAOYUN (NANJING) TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510324015.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Under the traditional server management mode, administrators need to manually manage grouping and regularly inspect, which is inefficient and prone to missed faults or delayed responses due to human operational errors. Especially in super-large-scale distributed systems, it is difficult to meet the requirements of efficiency, flexibility and reliability.

Method used

It provides an automatic server inspection method and system that integrates automated task arrangement, dynamic grouping management and intelligent patrol. It obtains user-set grouping rules and patrol rules through interactive interfaces, collects server parameter information and stores it in a distributed relational database, uses scheduling and execution tools to write a inspection task list, and patrol through python scripts, generates inspection reports and automatically issues alarm information.

Benefits of technology

It realizes efficient group management and automatic inspection of large-scale servers, improves management efficiency, reduces manual intervention, enhances the intelligence and reliability of inspections, promptly detects potential problems, and reduces operation and maintenance risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144408A_ABST
    Figure CN120144408A_ABST
Patent Text Reader

Abstract

The invention provides an automatic inspection method and system for a server. The method comprises the following steps: acquiring a grouping rule and an inspection rule set by a user on an interactive interface; according to the grouping rule and the inspection rule, parameter information of the target server is collected, and the collected parameter information is stored in a distributed relational database; compiling the collected parameter information into an inspection task list through a scheduling and execution tool; performing routing inspection on the collected parameter information through a scheduling and execution tool and a python script to obtain a routing inspection result; generating an inspection report according to the inspection result, and displaying the operation state, fault distribution and performance bottleneck of the target server through the inspection report; the inspection result is analyzed, and when the inspection result is abnormal, alarm information is automatically sent out; automatic task arrangement, dynamic grouping management and intelligent routing inspection are integrated, and efficient grouping management and automatic routing inspection of large-scale servers are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to an automatic inspection method and system for servers. Background Art

[0002] With the rapid expansion of the scale of data centers, the number and complexity of servers are increasing day by day. In the traditional operation and maintenance mode, administrators need to manually group and manage a large number of servers and conduct regular inspections. This method is not only inefficient but also prone to missed inspections or delayed responses to failures due to human operation errors. Especially when facing ultra-large-scale distributed systems, the traditional inspection method is difficult to meet the requirements of modern data centers for high efficiency, flexibility, and reliability.

[0003] Although automated operation and maintenance tools (such as Ansible, Puppet, Chef) have gradually become mainstream, there are still the following problems when these tools are applied to server grouping management and inspection:

[0004] 1) High difficulty in grouping management. Current tools mostly rely on static grouping methods and are difficult to dynamically adjust server groups. Especially when dealing with multiple types of servers (such as different hardware configurations, operating system versions, or network topologies), the flexibility of the grouping strategy is insufficient.

[0005] 2) Lack of intelligence in the inspection process. Traditional inspections mostly rely on fixed rules or scripts and cannot dynamically adapt to the needs of servers in different groups. At the same time, the inspection results can only provide a simple health status judgment and lack the ability of intelligent analysis and trend prediction.

[0006] 3) Limited scalability and compatibility. Current solutions are difficult to adapt to server hardware and management interfaces from different manufacturers. Especially when supporting hardware-level data collection and analysis (such as temperature, power consumption, hard disk health status, sensors, etc.), a large amount of additional development work is required. Summary of the Invention

[0007] In view of this, the purpose of the present invention is to provide an automatic inspection method and system for servers, which integrates automated task scheduling, dynamic grouping management, and intelligent inspection, and realizes efficient grouping management and automatic inspection of large-scale servers.

[0008] In the first aspect, an embodiment of the present invention provides an automatic inspection method for servers, and the method includes:

[0009] Obtain the grouping rules and inspection rules set by the user on the interaction interface;

[0010] Collect the parameter information of the target server according to the grouping rules and the inspection rules, and store the collected parameter information in a distributed relational database;

[0011] Write the collected parameter information into an inspection task list through a scheduling and execution tool;

[0012] Perform inspections on the collected parameter information through the scheduling and execution tool and a Python script to obtain inspection results;

[0013] Generate an inspection report from the inspection results, and display the operating status, fault distribution, and performance bottleneck of the target server through the inspection report;

[0014] Analyze the inspection results, and automatically send an alarm message when there are abnormalities in the inspection results.

[0015] Further, the target server includes a first cloud server and a second cloud server; collect the parameter information of the target server according to the grouping rule and the inspection rule, and store the collected parameter information in a distributed relational database, including:

[0016] Collect the hardware information and software information of the first cloud server through an information collection tool according to the grouping rule and the inspection rule to obtain first parameter information;

[0017] Collect the hardware information and software information of the second cloud server through a server management interface to obtain second parameter information;

[0018] Store the first parameter information and the second parameter information as dynamic library files, and save the first parameter information and the second parameter information to the distributed relational database.

[0019] Further, perform inspections on the collected parameter information through the scheduling and execution tool and a Python script to obtain inspection results, including:

[0020] Perform normalization processing on the collected parameter information through the scheduling and execution tool and a Python script to obtain processed parameter information;

[0021] Detect abnormal data from the processed parameter information through set rules and historical data;

[0022] When abnormal values are detected, record and mark them, and discard data that cannot be repaired.

[0023] Further, analyze the inspection results, and automatically send an alarm message when there are abnormalities in the inspection results, including:

[0024] Use Ansible to execute the nvidia-smi command to check whether the return is normal;

[0025] When an error message is returned, determine that the graphics card driver is abnormal and automatically send out the warning message.

[0026] Furthermore, the software information includes the CPU model, number of cores, memory capacity, hard disk size, and GPU card type; the hardware information includes sensor data, disk, network card, fan, and power supply.

[0027] In a second aspect, an embodiment of the present invention provides an automatic inspection system for a server, and the system includes:

[0028] An interaction interface module, configured to obtain grouping rules and inspection rules set by a user on an interaction interface;

[0029] A dynamic grouping and label management module, configured to collect parameter information of a target server according to the grouping rules and the inspection rules, and store the collected parameter information in a distributed relational database;

[0030] An inspection task management module, configured to compile the collected parameter information into an inspection task list through a scheduling and execution tool;

[0031] A data collection and processing module, configured to perform an inspection on the collected parameter information through the scheduling and execution tool and a Python script to obtain an inspection result;

[0032] A result visualization and report generation module, configured to generate an inspection report from the inspection result, and display the operating status, fault allocation, and performance bottleneck of the target server through the inspection report;

[0033] An inspection anomaly notification module, configured to analyze the inspection result, and automatically send out a warning message when the inspection result is abnormal.

[0034] Furthermore, the target server includes a first cloud server and a second cloud server; specifically, the dynamic grouping and label management module is configured to:

[0035] Collect the hardware information and software information of the first cloud server through an information collection tool according to the grouping rules and the inspection rules to obtain first parameter information;

[0036] Collect the hardware information and software information of the second cloud server through a server management interface to obtain second parameter information;

[0037] Store the first parameter information and the second parameter information as a dynamic library file, and save the first parameter information and the second parameter information to the distributed relational database.

[0038] Further, the data acquisition and processing module is specifically configured to:

[0039] Normalize the collected parameter information through the scheduling and execution tool and the Python script to obtain the processed parameter information;

[0040] Detect abnormal data from the processed parameter information through set rules and historical data;

[0041] When an abnormal value is detected, record and mark it, and discard the data that cannot be repaired.

[0042] In a third aspect, an embodiment of the present invention provides an electronic device, including a memory and a processor. A computer program that can run on the processor is stored on the memory. When the processor executes the computer program, the above-described method is implemented.

[0043] In a fourth aspect, an embodiment of the present invention provides a computer-readable medium having non-volatile program code executable by a processor, and the program code causes the processor to execute the above-described method.

[0044] An embodiment of the present invention provides an automatic inspection method and system for a server, including: obtaining grouping rules and inspection rules set by a user on an interaction interface; collecting parameter information of a target server according to the grouping rules and inspection rules, and storing the collected parameter information in a distributed relational database; writing the collected parameter information into an inspection task list through a scheduling and execution tool; performing an inspection on the collected parameter information through the scheduling and execution tool and a Python script to obtain an inspection result; generating an inspection report from the inspection result, and displaying the running status, fault allocation, and performance bottleneck of the target server through the inspection report; analyzing the inspection result, and automatically sending an alarm message when the inspection result is abnormal; integrating automated task orchestration, dynamic grouping management, and intelligent inspection, and realizing efficient grouping management and automatic inspection of a large-scale server.

[0045] Other features and advantages of the present invention will be described in the following specification, and in part, will be obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention are achieved and obtained by the structures specifically pointed out in the specification, claims, and drawings.

[0046] To make the above objectives, features, and advantages of the present invention more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, provides a detailed description as follows. Description of the Drawings

[0047] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0048] Figure 1 Flowchart of the automatic inspection method for the server provided in the first embodiment of the present invention;

[0049] Figure 2 Schematic diagram of the automatic inspection architecture of the server provided in the first embodiment of the present invention;

[0050] Figure 3 Schematic diagram of the CPU inspection of the server provided in the first embodiment of the present invention;

[0051] Figure 4 Schematic diagram of the inspection indicators provided in the first embodiment of the present invention;

[0052] Figure 5 Schematic diagram of the code taking an indicator as an example provided in the first embodiment of the present invention;

[0053] Figure 6 Schematic diagram of the inspection taking multiple indicators as an example provided in the first embodiment of the present invention;

[0054] Figure 7 Schematic diagram of the CPU and GPU indicator inspections provided in the first embodiment of the present invention;

[0055] Figure 8 Schematic diagram of the inspection results provided in the first embodiment of the present invention;

[0056] Figure 9 Schematic diagram of the nvidia graphics card driver exception check provided in the first embodiment of the present invention;

[0057] Figure 10 Schematic diagram of the driver exception situation provided in the first embodiment of the present invention;

[0058] Figure 11 Schematic diagram of another driver exception situation provided in the first embodiment of the present invention;

[0059] Figure 12 Schematic diagram of the alarm information notification provided in the first embodiment of the present invention;

[0060] Figure 13 Schematic diagram of the automatic inspection system of the server provided in the second embodiment of the present invention. Specific embodiments

[0061] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0062] The servers of this application have deep technical accumulation and practical experience in the fields of hardware and management. Including high-performance computing servers, storage servers, and edge computing servers, etc., which are widely used in the fields of cloud computing, big data, and artificial intelligence.

[0063] In actual use, the cloud servers of this application have the following remarkable features:

[0064] 1) High performance and high reliability. The cloud servers are equipped with self-developed hardware architectures and redundant designs to ensure stable operation under high loads and harsh environments.

[0065] 2) Intelligent management interfaces. The cloud servers support hardware status monitoring and management through self-developed scipmitool tools as well as IPMI, Redfish, and proprietary API interfaces, and can collect key data such as temperature, power consumption, and disk health status in real time.

[0066] 3) Modularity and scalability. The cloud servers are flexibly designed, support modular expansion, adapt to different business requirements, and can meet multi-scenario usage from small and medium-sized enterprises to ultra-large data centers.

[0067] The embodiments of the present invention provide an automatic inspection method and system for servers, aiming to maximize the hardware performance and management advantages of cloud servers, and improve the operation and maintenance efficiency of data centers through the following means:

[0068] 1) Dynamic grouping and tagging management. Utilize the hardware characteristics of cloud servers (such as the number of CPU cores, memory capacity, hard disk type, etc.) to achieve dynamic grouping and tagging management of servers, and formulate personalized inspection strategies for different groups.

[0069] 2) Automated and intelligent inspections. Based on the intelligent management interfaces of cloud servers (such as Redfish API, IPMI, SNMP, SSH protocols), collect hardware health data, and combine with the automation capabilities of Ansible to quickly complete inspection tasks. At the same time, the system introduces AI algorithms to perform intelligent analysis and fault prediction on the inspection results.

[0070] 3) Compatibility with multiple brands and multiple scenarios. In addition to supporting cloud servers, the system also has good compatibility, can integrate the management interfaces and inspection rules of multiple brands of servers, and meet the operation and maintenance requirements in mixed environments.

[0071] 4) Result visualization and data analysis. Generate real-time reports for the inspection tasks, showing key information including the server running status, hardware failure distribution, and performance bottlenecks, etc., providing intuitive management perspectives and scientific decision-making support for the operation and maintenance personnel.

[0072] In this embodiment, the cloud server can play a more efficient management efficiency in the actual operation and maintenance scenario, and at the same time provide a more intelligent data center operation and maintenance experience for customers, improving the market competitiveness.

[0073] This application aims to solve the problems of complex grouping, low inspection efficiency, and poor scalability in large-scale server management. The system realizes the automatic collection of server hardware and software information through dynamic grouping and tagging management, and dynamically updates the grouping according to the set rules; through the automated task scheduling and concurrent execution capabilities of Ansible, it completes the efficient inspection of server groups; combines AI algorithms to perform intelligent analysis and anomaly detection on the inspection data, providing accurate fault prediction and optimization suggestions for the operation and maintenance personnel.

[0074] At the same time, the present invention designs an architecture with high compatibility and scalability, supports the access of cloud servers and cloud servers of other brands, and allows users to customize inspection rules and task processes to adapt to the needs of different business scenarios. Through this application, the efficiency and intelligent level of server management and inspection are significantly improved, providing an important technical support for the efficient operation and maintenance of the data center.

[0075] For the convenience of understanding this embodiment, the embodiments of the present invention will be introduced in detail below.

[0076] Embodiment 1:

[0077] Figure 1 It is a flowchart of the automatic inspection method for the server provided in Embodiment 1 of the present invention.

[0078] Referring to Figure 1 , the method includes the following steps:

[0079] Step S101, obtain the grouping rules and inspection rules set by the user on the interaction interface;

[0080] Step S102, collect the parameter information of the target server according to the grouping rules and inspection rules, and store the collected parameter information in the distributed relational database;

[0081] Step S103, write the collected parameter information into an inspection task list through the scheduling and execution tool;

[0082] Step S104, perform inspections on the collected parameter information through the scheduling and execution tool and the python script to obtain inspection results;

[0083] Step S105: Generate an inspection report based on the inspection results, and display the operating status, fault allocation, and performance bottleneck of the target server through the inspection report.

[0084] Step S106: Analyze the inspection results, and automatically send an alarm message when there are abnormalities in the inspection results.

[0085] Specifically, referring to Figure 2 , this application mainly collects and stores the server information (such as operating system version, hardware model, network status) and hardware information (sensor data, disks, network cards, fans, power supplies, etc.) of the target server into a distributed relational database (GaussDB) based on the grouping rules and inspection rules set by the user and using a scheduling and execution tool (ansible-playbook). The system automatically analyzes the information collected during the inspection and dynamically updates the grouping and fault notification reports according to the set rules.

[0086] The core modules included in this platform are described as follows:

[0087] Web UI: It provides an interactive interface for configuration, including server configuration, grouping list, inspection item list, addition and modification of inspection lists, and display of inspection results. It allows users to customize the grouping and inspection of servers, and users can access it through a browser. As Figure 3 shown, perform a CPU inspection on a cloud server of type R5215A12. The inspection metrics are as Figure 4 shown.

[0088] Dynamic grouping and label management module: After configuring the server list through the Web UI, for cloud servers, the hardware information and software information (such as CPU model, number of cores, memory capacity, hard disk size, GPU card type, etc.) of the servers can be uniformly collected and parsed through the scipmitool tool. For cloud servers of other manufacturers, they can be collected and parsed through server management interfaces (such as IPMI, Redfish, SSH, SNMP), stored as dynamic library files, and at the same time, the information is also saved to the distributed relational database. Taking the number of CPU cores, memory capacity, and hard disk size as metrics, the example code is as Figure 5 shown.

[0089] The inspection task management module is responsible for task definition, distribution, execution, and status tracking. The inspection is divided into manual inspection and scheduled automatic inspection. The following is an example of adding an automatic inspection task to inspect multiple metrics of the CPU and GPU of a cloud server of type R5215A12. Specifically, it is as Figure 6 shown.

[0090] After the user configures the inspection task in the WEB UI, the system backend will generate an inspection task list based on the scheduling and execution tool (ansible-Playbook), and use a Python script in combination with the scheduling and execution tool to efficiently complete the inspection.

[0091] Further, the target servers include a first cloud server and a second cloud server; step S102 includes the following steps:

[0092] Step S201, according to the grouping rule and the inspection rule, collect the hardware information and software information of the first cloud server through the information collection tool to obtain the first parameter information;

[0093] Step S202, collect the hardware information and software information of the second cloud server through the server management interface to obtain the second parameter information;

[0094] Step S203, store the first parameter information and the second parameter information as a dynamic library file, and save the first parameter information and the second parameter information to a distributed relational database. Among them, the second cloud server is a server of other brands.

[0095] Further, step S104 includes the following steps:

[0096] Step S301, normalize the collected parameter information through the scheduling and execution tool and the Python script to obtain the processed parameter information;

[0097] Step S302, detect abnormal data in the processed parameter information through the set rules and historical data;

[0098] Step S303, when an abnormal value is detected, record and mark it, and discard the data that cannot be repaired.

[0099] Specifically, in the data collection and processing module, the unified collection of hardware data is carried out through the information collection tool (scipmitool). For servers of other brands, hardware data such as temperature, power consumption, and disk health status, as well as operating system performance data such as CPU, memory, and disk IO, are collected through multi-protocol interfaces (such as IPMI, Redfish, SNMP, SSH). The collected data is discarded (such as normalization, abnormal value detection), and the results are saved to GaussDB to prepare for subsequent analysis and display.

[0100] Normalize the collected parameter information. Since the data formats and units of different protocols may be different, the data is converted into a unified standard format: temperature: °C; power consumption: watt (W); time: unified into UTC format.

[0101] Abnormal data detection and processing. Using statistics to detect outliers:

[0102] 1) Based on rules: Set a reasonable range for the hardware. For example: Temperature range: 0 - 80 °C; CPU usage range: 0% - 100%.

[0103] 2) Based on historical data: Compare the current value with the historical trend to detect sudden increases or abnormal changes.

[0104] Processing methods for detected outliers:

[0105] 1) Record and mark (such as is_abnormal = true); 2) Discard data that cannot be repaired.

[0106] Furthermore, step S106 includes the following steps:

[0107] Step S401, execute the nvidia-smi command through ansible to check whether the return is normal;

[0108] Step S402, when an error message is returned, determine that the graphics card driver is abnormal and automatically send an alarm message.

[0109] Furthermore, software information includes CPU model, number of cores, memory capacity, hard disk size, and GPU card type; hardware information includes sensor data, disk, network card, fan, and power supply.

[0110] Result visualization and report generation module. Visualize the inspection results and generate an interactive inspection report. Develop the front-end page using the Web front-end technology Vue.js to display the server running status, fault allocation, performance bottlenecks, etc. Automatically generate an inspection report and support export to Excel format. Support historical data query and comparative analysis based on time and grouping. The following shows the inspection results and inspection result details: Taking the CPU and GPU metrics inspection of the R5215A12 type cloud server as an example, as Figure 7 shown. The inspection results are as Figure 8 shown.

[0111] Inspection anomaly notification module: This module continuously analyzes and processes based on the inspection results. When there are anomalies in the inspection results, it will automatically send a warning notification to the operation and maintenance personnel. Taking the inspection of the nvidia GPU card as an example, refer to Figure 9 .

[0112] nvidia graphics card driver anomaly check. Issue an inspection task through the inspection task management module, and execute the nvidia-smi command through ansible to check whether the return is normal.

[0113] If the following error is returned, it indicates that the graphics card driver is abnormal, and the system needs to report an alarm to the operation and maintenance personnel. For the first case of driver abnormality, refer to Figure 10 , for the second case of driver abnormality, refer to Figure 11 .

[0114] The system supports notifying the operation and maintenance personnel via alarm configuration emails and text messages. Refer to Figure 12 .

[0115] Embodiment 2:

[0116] Figure 13 It is a schematic diagram of the automatic inspection system of the server provided by the second embodiment of the present invention.

[0117] Refer to Figure 13 , the system includes:

[0118] An interaction interface module, which is used to obtain the grouping rules and inspection rules set by the user on the interaction interface;

[0119] A dynamic grouping and label management module, which is used to collect the parameter information of the target server according to the grouping rules and inspection rules, and store the collected parameter information in a distributed relational database;

[0120] An inspection task management module, which is used to compile the collected parameter information into an inspection task list through a scheduling and execution tool;

[0121] A data collection and processing module, which is used to inspect the collected parameter information through a scheduling and execution tool and a python script to obtain an inspection result;

[0122] A result visualization and report generation module, which is used to generate an inspection report from the inspection result, and display the running status, fault allocation, and performance bottleneck of the target server through the inspection report;

[0123] An inspection exception notification module, which is used to analyze the inspection result, and automatically send an alarm message when there is an exception in the inspection result.

[0124] Specifically, based on the inspection tasks written by the scheduling and execution tool, inspection strategies corresponding to different groups are defined (such as hardware inspection, log analysis, resource utilization statistics, etc.)

[0125] The cloud server collects hardware data such as temperature, power consumption, and disk health status, as well as operating system performance data such as CPU, memory, and disk I / O through a self-developed information collection tool (scipmitool), and other brand servers collect data through multi-protocol interfaces (such as IPMI, Redfish, SNMP, SSH).

[0126] Use self-developed algorithms, combined with inspection historical data and real-time data, to analyze the health status of cloud servers. Generate an exception report and attach potential fault points (such as hard disk failures, fan anomalies, memory overflows, etc.) for fault notification.

[0127] Information storage: The system will collect system performance and hardware information of each target device, including but not limited to GPU, CPU, memory, disk, temperature, fan, power sensor, network, process, etc., and store this data in the GaussDB database of the centralized management server for management and analysis.

[0128] Log recording: All operation logs, error messages, etc. of grouping and inspection will be uniformly saved in the database for auditing and traceability.

[0129] Furthermore, the target server includes a first cloud server and a second cloud server; the dynamic grouping and label management module is specifically used for:

[0130] According to the grouping rules and inspection rules, collect the hardware information and software information of the first cloud server through the information collection tool to obtain the first parameter information;

[0131] Collect the hardware information and software information of the second cloud server through the server management interface to obtain the second parameter information;

[0132] Store the first parameter information and the second parameter information as dynamic library files, and save the first parameter information and the second parameter information to the distributed relational database.

[0133] Furthermore, the data collection and processing module is specifically used for:

[0134] Normalize the collected parameter information through the scheduling and execution tool and the python script to obtain the processed parameter information;

[0135] Detect abnormal data in the processed parameter information through the set rules and historical data;

[0136] When an abnormal value is detected, record and mark it, and discard the data that cannot be repaired.

[0137] The beneficial effects of this application are:

[0138] 1) Improve server management efficiency. Implement dynamic grouping and labeling management of servers, support flexible adjustment of grouping rules, reduce manual intervention, and improve management efficiency.

[0139] 2) Improve the level of inspection automation. Based on ansible's automated task scheduling and concurrent execution, quickly complete the inspection of large-scale servers, significantly reducing the time cost.

[0140] 3) Enhance the intelligent capabilities of patrol inspections. Combine AI algorithms to perform anomaly detection and fault prediction on patrol inspection data, promptly identify potential problems, and reduce operation and maintenance risks.

[0141] 4) Optimize the utilization of operation and maintenance resources. Through intelligent analysis and visual reports, operation and maintenance personnel can accurately locate problem servers, improving resource utilization efficiency and decision-making effectiveness.

[0142] 5) Strong compatibility and scalability. Support the access of cloud servers and other multi-brand cloud servers, provide a plug-in design, and meet the patrol inspection requirements of different scenarios.

[0143] This application effectively solves the problems of low efficiency, poor flexibility, and insufficient scalability in traditional server operation and maintenance, providing a comprehensive solution for the intelligent and automated operation and maintenance of data centers.

[0144] An embodiment of the present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the automatic patrol inspection method of the server provided in the above embodiment are implemented.

[0145] An embodiment of the present invention also provides a computer-readable medium having non-volatile program code executable by a processor. A computer program is stored on the computer-readable medium, and when the computer program is run by the processor, the steps of the automatic patrol inspection method of the server in the above embodiment are executed.

[0146] The computer program product provided by the embodiment of the present invention includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the method described in the foregoing method embodiments. For specific implementation, reference can be made to the method embodiments, and details are not described herein again.

[0147] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems and devices described above can refer to the corresponding processes in the foregoing method embodiments, and details are not described herein again.

[0148] In addition, in the description of the embodiments of the present invention, unless otherwise clearly defined and limited, the terms "installation", "connection", and "connection" shall be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0149] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0150] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention. In addition, the terms "first", "second", "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.

[0151] Finally, it should be noted that the above-mentioned embodiments are only specific embodiments of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting them. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions described in the foregoing embodiments, or can easily think of changes, or make equivalent replacements for some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for automatic inspection of a server, characterized in that: The method comprises: Obtain the grouping rules and inspection rules set by the user on the interactive interface; According to the grouping rule and the inspection rule, parameter information of the target server is collected, and the collected parameter information is stored in a distributed relational database; The collected parameter information is compiled into an inspection task list through a scheduling and execution tool; The collected parameter information is inspected by using the scheduling and execution tool and the python script to obtain an inspection result; Generate an inspection report based on the inspection result, and display the running status, fault distribution and performance bottleneck of the target server through the inspection report; The inspection results are analyzed, and when there are abnormalities in the inspection results, an alarm message is automatically issued.

2. The automatic inspection method of a server according to claim 1, characterized in that: The target server includes a first cloud server and a second cloud server; according to the grouping rule and the inspection rule, parameter information of the target server is collected, and the collected parameter information is stored in a distributed relational database, including: According to the grouping rule and the inspection rule, the hardware information and software information of the first cloud server are collected by an information collection tool to obtain first parameter information; Collecting hardware information and software information of the second cloud server through a server management interface to obtain second parameter information; The first parameter information and the second parameter information are stored as a dynamic library file, and the first parameter information and the second parameter information are saved in the distributed relational database.

3. The automatic inspection method of a server according to claim 1, characterized in that: The collected parameter information is inspected by the scheduling and execution tool and the python script to obtain inspection results, including: The collected parameter information is normalized by the scheduling and execution tool and the python script to obtain processed parameter information; Perform abnormal data detection on the processed parameter information by setting rules and historical data; When outliers are detected, they are recorded and flagged, and data that cannot be corrected is discarded.

4. The automatic inspection method of a server according to claim 1, characterized in that: The inspection results are analyzed, and when the inspection results are abnormal, alarm information is automatically issued, including: Run the nvidia-smi command through ansible to check whether the response is normal; When an error message is returned, it is determined that the graphics card driver is abnormal, and the warning message is automatically issued.

5. The automatic inspection method of a server according to claim 2, characterized in that: The software information includes CPU model, number of cores, memory capacity, hard disk size and GPU card type; the hardware information includes sensor data, disk, network card, fan and power supply.

6. An automatic inspection system for a server, characterized in that: The system comprises: The interactive interface module is used to obtain the grouping rules and inspection rules set by the user on the interactive interface; A dynamic grouping and label management module, used to collect parameter information of the target server according to the grouping rule and the inspection rule, and store the collected parameter information in a distributed relational database; An inspection task management module is used to compile the collected parameter information into an inspection task list through a scheduling and execution tool; A data collection and processing module is used to inspect the collected parameter information through the scheduling and execution tool and the python script to obtain the inspection result; A result visualization and report generation module is used to generate an inspection report from the inspection results, and display the operating status, fault distribution and performance bottleneck of the target server through the inspection report; The inspection abnormality notification module is used to analyze the inspection results and automatically issue an alarm message when the inspection results are abnormal.

7. The automatic inspection system for servers according to claim 6, characterized in that: The target server includes a first cloud server and a second cloud server; the dynamic grouping and tag management module is specifically used for: According to the grouping rule and the inspection rule, the hardware information and software information of the first cloud server are collected by an information collection tool to obtain first parameter information; Collecting hardware information and software information of the second cloud server through a server management interface to obtain second parameter information; The first parameter information and the second parameter information are stored as a dynamic library file, and the first parameter information and the second parameter information are saved in the distributed relational database.

8. The automatic inspection system for servers according to claim 6, characterized in that: The data acquisition and processing module is specifically used for: The collected parameter information is normalized by the scheduling and execution tool and the python script to obtain processed parameter information; Perform abnormal data detection on the processed parameter information by setting rules and historical data; When outliers are detected, they are recorded and flagged, and data that cannot be corrected is discarded.

9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 5 is implemented.

10. A computer readable medium having a non-volatile program code executable by a processor, characterized in that: The program code enables the processor to execute the method according to any one of claims 1 to 5.