Automatic inspection method and device and readable storage medium
By using automated scripts and localized NLP models, the problems of low efficiency, insufficient intelligence, and security risks in network device inspection have been solved, achieving efficient and secure intelligent operation and maintenance management.
Patent Information
- Application Number
- CN202510990241.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-10-31
AI Technical Summary
In existing technologies, network equipment inspection relies on manual operation, which is inefficient, has insufficient intelligence in log analysis, lacks intelligent assistance, and poses security risks in data flow.
The system uses pre-written automated scripts from terminal emulation software to log into multiple devices, acquire and store raw logs, perform preprocessing and extract key indicators, use a fine-tuned NLP model combined with historical data to perform anomaly detection and root cause diagnosis, generate operation and maintenance suggestions, and complete all processing locally.
It has achieved efficient and automated inspection, improved the level of intelligent log analysis, reduced manual operation, ensured data security, generated structured analysis reports, and supported intelligent operation and maintenance decision-making.
Smart Images

Figure CN120880873A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of operation and maintenance management technology, and in particular to an automatic inspection method, device and readable storage medium. Background Technology
[0002] As enterprise IT (Information Technology) infrastructure expands, routine inspections of network equipment (such as servers, routers, and switches) have become a core task of operations and maintenance management. Traditional inspection methods mainly rely on manual operation or basic script tools, which have the following drawbacks:
[0003] 1) Manual inspection is inefficient and has poor fault tolerance: Maintenance personnel need to log in to each device one by one (e.g., manually enter commands through SecureCRT), which is time-consuming and labor-intensive, and critical devices are easily missed due to operational errors. For large-scale clusters, manual inspection is difficult to meet real-time requirements, and delays in fault response may cause business interruption.
[0004] 2) Insufficient intelligence in log analysis: Although existing script tools can collect logs, they lack structured analysis capabilities. The raw log data is mixed and redundant, requiring manual screening of abnormal fields and generation of statistical reports, resulting in low analysis efficiency.
[0005] 3) Decision-making relies on experience and lacks intelligent assistance: Anomaly diagnosis and handling suggestions heavily depend on the experience of operations and maintenance personnel. Traditional solutions cannot combine historical data or business scenarios to generate adaptive decisions (such as predictive maintenance suggestions). Existing technologies rely on cloud-based AI (Artificial Intelligence) analysis tools. While public cloud APIs (Application Programming Interfaces) can partially address this, they lack the necessary support.
[0006] This approach offers a solution to the problem, but is difficult to deploy directly due to enterprise intranet data security requirements.
[0007] 4) Data transfer poses security risks: If cloud-based AI is used to process logs, sensitive operational data needs to be uploaded to a third-party server, which may violate the company's data privacy regulations. On the other hand, traditional fully localized solutions cannot achieve intelligent analysis, creating a conflict between security and functionality. Summary of the Invention
[0008] The technical problem to be solved by this invention is to address the above-mentioned shortcomings of the prior art by providing an automatic inspection method, device, and readable storage medium, so as to solve the problem that existing inspection technologies mainly rely on manual operation or basic script tools, resulting in low efficiency of manual inspection.
[0009] The log analysis suffers from insufficient intelligence, lack of intelligent assistance, and security risks in data flow.
[0010] In a first aspect, the present invention provides an automatic inspection method, the method comprising:
[0011] The system automatically logs into multiple devices and executes inspection commands using a pre-written automated login script for terminal emulation software.
[0012] Obtain the raw logs output by multiple devices after executing the inspection command, and store the raw logs in a local database;
[0013] The stored raw logs are preprocessed, and key indicators are extracted from the preprocessed raw logs using preset indicator extraction rules.
[0014] Anomaly detection is performed on the extracted key indicators to obtain the abnormal indicators and the cross-device correlation anomaly results of the corresponding devices;
[0015] Based on the aforementioned abnormal indicators and the cross-device correlation abnormal results of the corresponding devices, the root cause diagnosis results are output using a fine-tuned natural language processing (NLP) model combined with the historical time series data of the corresponding devices. The fine-tuned NLP model is deployed on a local server.
[0016] Based on the root cause diagnosis results and the pre-built operation and maintenance knowledge graph, corresponding operation and maintenance suggestions are obtained, and a structured analysis report is output.
[0017] Furthermore, the terminal emulation software is SecureCRT, and the automatic login to multiple devices and execution of inspection commands via a pre-written automated login script of the terminal emulation software specifically includes:
[0018] The system uses a thread pool to concurrently log in to multiple devices using a pre-written SecureCRT automated login script and pre-encrypted and stored usernames and passwords for each device.
[0019] Inspection commands are sent to multiple devices in a preset order.
[0020] Furthermore, the step of obtaining the raw logs output by the multiple devices after executing the inspection command, and storing the raw logs in a local database, specifically includes:
[0021] The system captures the raw logs output by multiple devices after executing the inspection command in real time, removes control characters, and retains plain text data.
[0022] The raw log data of plain text is categorized and stored in a local relational database or time-series database.
[0023] Furthermore, the preprocessing of the stored raw logs specifically includes:
[0024] The stored raw logs are formatted and cleaned to remove garbled characters and duplicate spaces.
[0025] The timestamps in the original logs from multiple devices, after being format-cleaned, are uniformly converted to Coordinated Universal Time (UTC) format.
[0026] Furthermore, the step of performing anomaly detection on the extracted key indicators to obtain abnormal indicators and corresponding cross-device correlation anomaly results for the devices specifically includes:
[0027] The extracted key indicators are then normalized.
[0028] Based on the preset rule engine, the threshold evaluation of the normalized key indicators is performed to obtain the first abnormal indicator that triggers the static threshold alarm.
[0029] By calculating the historical data baseline using a pre-set statistical model, the second abnormal indicator that deviates from the normal range among the key indicators after normalization is identified.
[0030] Based on the devices corresponding to the first and second anomaly indicators, cross-device association anomalies are detected using a pre-built graph database-based device topology relationship, and corresponding results are obtained.
[0031] Furthermore, before outputting the root cause diagnosis result based on the abnormal indicators and the cross-device correlation anomaly results of the corresponding devices using a fine-tuned Natural Language Processing (NLP) model combined with the historical time-series data of the corresponding devices, the method further includes:
[0032] The fine-tuned NLP model is packaged into a Docker service and deployed on a local server.
[0033] Furthermore, the step of obtaining corresponding operation and maintenance suggestions based on the root cause diagnosis results and the pre-constructed operation and maintenance knowledge graph, and outputting a structured analysis report, specifically includes:
[0034] Based on a pre-built operation and maintenance knowledge graph, historical faults similar to the root cause diagnosis results and corresponding solutions are matched.
[0035] Based on the solution, a corresponding repair strategy is determined, and the repair strategy is converted into a natural language description to obtain corresponding operation and maintenance suggestions;
[0036] Based on the aforementioned maintenance recommendations, a corresponding structured analysis report is generated using a template engine.
[0037] Secondly, the present invention provides an automatic inspection device, the device comprising:
[0038] The automatic login inspection module is used to automatically log in to multiple devices and execute inspection commands through a pre-written automated login script of terminal emulation software.
[0039] The raw log storage module is connected to the automatic login inspection module and is used to obtain the raw logs output by multiple devices after executing the inspection command, and store the raw logs in a local database.
[0040] The key indicator extraction module is connected to the original log storage module and is used to preprocess the stored original logs and extract key indicators from the preprocessed original logs using preset indicator extraction rules.
[0041] An anomaly detection module, connected to the key indicator extraction module, is used to perform anomaly detection on the extracted key indicators and obtain the anomaly indicators and the cross-device associated anomaly results of the corresponding devices.
[0042] The root cause diagnosis module, connected to the anomaly detection module, is used to output the root cause diagnosis result based on the anomaly indicators and the cross-device correlation anomaly results of the corresponding devices, using a fine-tuned natural language processing (NLP) model combined with the historical time series data of the corresponding devices. The fine-tuned NLP model is deployed on a local server.
[0043] The operation and maintenance suggestion module is connected to the root cause diagnosis module. It is used to obtain corresponding operation and maintenance suggestions based on the root cause diagnosis results and the pre-built operation and maintenance knowledge graph, and output a structured analysis report.
[0044] Thirdly, the present invention provides an automatic inspection device, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to implement the automatic inspection method described in the first aspect above.
[0045] Fourthly, the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the automatic inspection method described in the first aspect.
[0046] This invention provides an automated inspection method, apparatus, and readable storage medium. First, multiple devices are automatically logged in using a pre-written automated login script from terminal emulation software, and inspection commands are executed. Then, the raw logs output by the multiple devices after executing the inspection commands are acquired and stored in a local database. Next, the stored raw logs are preprocessed, and key indicators are extracted from the preprocessed logs using preset indicator extraction rules. Anomaly detection is performed on the extracted key indicators to obtain abnormal indicators and cross-device correlation anomaly results for corresponding devices. Based on the abnormal indicators and the cross-device correlation anomaly results, a fine-tuned Natural Language Processing (NLP) model combined with the historical time-series data of the corresponding devices is used to output root cause diagnosis results, wherein the fine-tuned NLP model is deployed on a local server. Finally, based on the root cause diagnosis results and a pre-constructed operation and maintenance knowledge graph, corresponding operation and maintenance suggestions are obtained, and a structured analysis report is output. This invention achieves automated inspection of multiple devices through an automated login script, improving inspection efficiency and reducing manual operation. Meanwhile, by extracting key indicators and utilizing anomaly detection technology and a fine-tuned NLP model combined with historical data to output root cause diagnosis results, the intelligence level of log analysis has been significantly improved. Furthermore, generating operational suggestions based on an operational knowledge graph enhances intelligent auxiliary functions. By storing raw logs in a local database and deploying the fine-tuned NLP model on a local server, data security across the entire chain is ensured, achieving automated and intelligent operational management. This addresses the shortcomings of existing inspection technologies, which primarily rely on manual operation or basic script tools, resulting in low efficiency in manual inspections.
[0047] The log analysis suffers from insufficient intelligence, lack of intelligent assistance, and security risks in data flow. Attached Figure Description
[0048] Figure 1 This is a flowchart of an automatic inspection method according to Embodiment 1 of the present invention;
[0049] Figure 2 This is an architecture diagram of the automatic inspection system according to an embodiment of the present invention;
[0050] Figure 3 This is a schematic diagram of the structure of an automatic inspection device according to Embodiment 2 of the present invention;
[0051] Figure 4 This is a schematic diagram of an automatic inspection device according to Embodiment 3 of the present invention. Detailed Implementation
[0052] To enable those skilled in the art to better understand the technical solution of the present invention, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.
[0053] It is understood that the specific embodiments and accompanying drawings described herein are merely for explaining the invention and are not intended to limit the invention.
[0054] It is understood that, without conflict, the various embodiments and features in the embodiments of the present invention can be combined with each other.
[0055] It is understood that, for ease of description, only the parts related to the present invention are shown in the accompanying drawings, while the parts unrelated to the present invention are not shown in the drawings.
[0056] It is understood that each unit or module involved in the embodiments of the present invention may correspond to only one entity structure, or may be composed of multiple entity structures, or multiple units or modules may be integrated into one entity structure.
[0057] It is understood that the terms "first," "second," etc., in the embodiments of the present invention are used to distinguish different objects or to distinguish different treatments of the same object, rather than to describe a specific order of objects.
[0058] It is understood that, without conflict, the functions and steps marked in the flowcharts and block diagrams of this invention may occur in a different order than that marked in the accompanying drawings.
[0059] It is understood that the flowcharts and block diagrams of this invention illustrate the possible architecture, functions, and operations of systems, apparatuses, devices, and methods according to various embodiments of this invention. Each block in the flowchart or block diagram may represent a unit, module, program segment, or code, containing executable instructions for implementing the specified function. Furthermore, each block or combination of blocks in the block diagram and flowchart can be implemented using a hardware-based system to achieve the specified function, or using a combination of hardware and computer instructions.
[0060] It is understood that the units and modules involved in the embodiments of the present invention can be implemented by software or by hardware. For example, the units and modules can be located in a processor.
[0061] Example 1:
[0062] This embodiment provides an automatic inspection method, such as Figure 1 As shown, the method includes:
[0063] Step S101: Automatically log in to multiple devices and execute inspection commands using a pre-written automated login script from the terminal emulation software;
[0064] In this embodiment, the terminal emulation software can be tools that support SSH (Secure Shell) / Telnet (Teletype Network) protocols, such as SecureCRT, PuTTY, and Xshell. Multiple devices refer to the devices that need to be inspected, including but not limited to servers, routers, and switches.
[0065] Optionally, the terminal emulation software is SecureCRT, and the automatic login to multiple devices and execution of inspection commands via a pre-written automated login script of the terminal emulation software specifically includes:
[0066] The system uses a thread pool to concurrently log in to multiple devices using a pre-written SecureCRT automated login script and pre-encrypted and stored usernames and passwords for each device.
[0067] Inspection commands are sent to multiple devices in a preset order.
[0068] In this embodiment, the terminal emulation software is preferably SecureCRT. Using the Python scripting language supported by SecureCRT, an automated login script is written to connect to multiple devices via SSH / Telnet protocol. The device's IP address (Internet Protocol), port, username, and password can be stored in a local encrypted configuration file. The script can dynamically read these configurations and adapt to different device authentication methods, such as RSA (Rivest-Shamir-Adleman, an asymmetric encryption algorithm) keys or two-factor authentication. Furthermore, concurrent login from multiple devices is achieved through thread pool technology, and a timeout retry mechanism is set (e.g., 3 retries, each with a 10-second interval) to avoid single points of failure causing overall process interruption.
[0069] In this embodiment, commonly used inspection commands (such as show cpu usage, display interface brief) can be preset, and the syntax can be automatically switched according to the device type. After successful login, the script sends inspection commands to the device in a preset order.
[0070] Step S102: Obtain the raw logs output by the multiple devices after executing the inspection command, and store the raw logs in the local database.
[0071] In this embodiment, command response status (such as Success / Error) can be matched in real time using regular expressions. If a command execution failure is detected (e.g., connection interruption or syntax error), the system will automatically record an error log and can choose to trigger a retry operation or mark the device as "requiring manual intervention". The original log includes all output of the command execution, whether successful or not.
[0072] Optionally, obtaining the raw logs output by the multiple devices after executing the inspection command, and storing the raw logs in a local database, specifically includes:
[0073] The system captures the raw logs output by multiple devices after executing the inspection command in real time, removes control characters, and retains plain text data.
[0074] The raw log data of plain text is categorized and stored in a local relational database or time-series database.
[0075] In this embodiment, control characters, such as ANSI color codes, are removed, leaving only the plain text data in the original logs. The original logs, containing only plain text data, are then categorized by device IP, timestamp, and command type, and stored in a local relational database (such as MySQL) or time-series database (such as InfluxDB). An index is created to accelerate subsequent queries. The command type can be categorized by device metrics such as CPU (central processing unit) utilization, memory usage, and port status.
[0076] Step S103: Preprocess the stored raw logs and extract key indicators from the preprocessed raw logs using preset indicator extraction rules.
[0077] In this embodiment, metric extraction rules can be defined for each type of device. For example, for Huawei devices, the CPU utilization extraction rule can be defined as a regular expression: CPU usage.*? (\d+)%). By calling Python's re library or open-source parsing tools (such as Grok), key metrics in the logs, such as memory usage and port traffic, can be matched.
[0078] Optionally, the preprocessing of the stored raw logs specifically includes:
[0079] The stored raw logs are formatted and cleaned to remove garbled characters and duplicate spaces.
[0080] The timestamps in the original logs from multiple devices, after being format-cleaned, are uniformly converted to UTC (Coordinated Universal Time) format.
[0081] In this embodiment, regular expressions can be used to remove irrelevant characters (such as garbled characters and repeated spaces) from the original logs, and then the timestamps of different devices can be uniformly converted to UTC format to solve the problem of inconsistent time zones. In addition, irrelevant log lines (such as debugging information) can be discarded based on a preset rule base (such as whitelist keywords).
[0082] Step S104: Perform anomaly detection on the extracted key indicators to obtain the abnormal indicators and the cross-device correlation anomaly results of the corresponding devices.
[0083] In this embodiment, abnormal indicators can be obtained based on threshold alarms and dynamic baseline detection, and cross-device related abnormal results of corresponding devices can be obtained through correlation analysis.
[0084] Optionally, the step of performing anomaly detection on the extracted key indicators to obtain abnormal indicators and corresponding cross-device correlation anomaly results for the devices specifically includes:
[0085] The extracted key indicators are then normalized.
[0086] Based on the preset rule engine, the threshold evaluation of the normalized key indicators is performed to obtain the first abnormal indicator that triggers the static threshold alarm.
[0087] By calculating the historical data baseline using a pre-set statistical model, the second abnormal indicator that deviates from the normal range among the key indicators after normalization is identified.
[0088] Based on the devices corresponding to the first and second anomaly indicators, cross-device association anomalies are detected using a pre-built graph database-based device topology relationship, and corresponding results are obtained.
[0089] In this embodiment, the extracted key indicators are first normalized (unified to standardized units, such as MB and Gbps); then, based on a preset rule engine (such as Drools), the normalized key indicators are subjected to threshold evaluation (automatically comparing the normalized key indicator data with preset static threshold rules). When a key indicator meets the static threshold condition (such as CPU > 90% for 5 minutes), the first abnormal indicator that triggers an alarm is determined; then, a historical data baseline is calculated using a statistical model (such as Z-Score or moving average) to identify the second abnormal indicator that deviates from the normal range; finally, based on the devices corresponding to the above two types of abnormal indicators, the device topology relationship pre-built in a graph database (such as Neo4j) is used to detect cross-device correlation anomalies (such as packet loss at a switch port causing increased latency in downstream servers).
[0090] Step S105: Based on the abnormal indicators and the cross-device correlation abnormal results of the corresponding devices, the root cause diagnosis results are output using the fine-tuned NLP (Natural Language Processing) model combined with the historical time series data of the corresponding devices. The fine-tuned NLP model is deployed on a local server.
[0091] In this embodiment, the fine-tuned NLP model can be a DeepSeek model based on BERT (Bidirectional Encoder Representations from Transformers). Anomaly indicators, cross-device correlation anomaly results for corresponding devices, and historical time-series data of the corresponding devices are taken as input, preprocessed, and integrated into the input vector of the DeepSeek model. The DeepSeek model performs inference by calling a local inference interface and outputs root cause diagnosis results.
[0092] Optionally, before outputting the root cause diagnosis result based on the abnormal indicators and the cross-device correlation abnormal results of the corresponding devices using a fine-tuned Natural Language Processing (NLP) model combined with the historical time-series data of the corresponding devices, the method further includes:
[0093] The fine-tuned NLP model is packaged into a Docker service and deployed on a local server.
[0094] In this embodiment, the NLP model can be fine-tuned using tagged historical fault data or normal state baseline data. The fine-tuned model is packaged as a Docker service and deployed on a local server to ensure that data processing and model inference are completed locally, avoiding leakage of sensitive data and providing rapid response.
[0095] Step S106: Based on the root cause diagnosis results and the pre-built operation and maintenance knowledge graph, obtain the corresponding operation and maintenance suggestions and output a structured analysis report.
[0096] In this embodiment, the fine-tuned NLP model (such as the DeepSeek model) can output a corresponding confidence score in addition to the root cause diagnosis result. If the confidence score is greater than a set threshold, corresponding operation and maintenance suggestions are generated based on the root cause diagnosis result and the pre-built operation and maintenance knowledge graph, and a structured analysis report is output. If the confidence score is lower than the set threshold, the root cause diagnosis result can be marked as "requiring manual intervention" for further verification.
[0097] Optionally, the step of obtaining corresponding operation and maintenance suggestions based on the root cause diagnosis results and the pre-built operation and maintenance knowledge graph, and outputting a structured analysis report, specifically includes:
[0098] Based on a pre-built operation and maintenance knowledge graph, historical faults similar to the root cause diagnosis results and corresponding solutions are matched.
[0099] Based on the solution, a corresponding repair strategy is determined, and the repair strategy is converted into a natural language description to obtain corresponding operation and maintenance suggestions;
[0100] Based on the aforementioned maintenance recommendations, a corresponding structured analysis report is generated using a template engine.
[0101] In this embodiment, an operations and maintenance knowledge graph is used to query historical fault cases and their solutions that are similar to the root cause diagnosis results. A repair strategy matching the solution (e.g., "reboot the port") is selected from the strategy template library. A template engine (e.g., Jinja2) or a lightweight NLP model (e.g., T5) is used to convert the repair strategy into a natural language description (e.g., "It is recommended to reboot the GigabitEthernet0 / 1 port after 17:00 today"), obtaining corresponding operations and maintenance recommendations. Finally, a structured analysis report containing visual charts is generated through the template engine, supporting export in multiple formats such as PDF (Portable Document Format) and HTML (Hyper Text Markup Language). This structured analysis report includes the corresponding operations and maintenance recommendations, the source of the abnormal data, etc.
[0102] In this embodiment, after generating the structured analysis report, it can be notified to operations and maintenance personnel. For example, it can be integrated with enterprise IM (Instant Messaging) bots (such as DingTalk and Lark) to send the report link and summary to designated groups via webhooks. Alternatively, it can use the SMTP (Simple Mail Transfer Protocol) protocol to send encrypted emails with the full report and execution suggestion tickets attached. If an emergency fault is detected (such as a core device outage), a telephone alarm is automatically triggered (e.g., via the Twilio API).
[0103] It should be noted that the automatic inspection method provided by this invention achieves safe and controllable end-to-end automated intelligent inspection, and its specific beneficial effects include:
[0104] 1) Achieve full-process automation: By integrating the SecureCRT command control module, it can automatically log in to multiple devices to perform inspection tasks, eliminating delays and errors caused by manual operation, and supporting concurrent inspections and real-time status monitoring of large-scale clusters.
[0105] 2) Improve the intelligence level of log analysis: Design a lightweight data processing module to perform structured parsing of raw logs (such as regular expression matching of key indicators and time series data aggregation), and combine statistical models and rule engines to identify abnormal patterns (such as memory leak trends and port conflict alarms) and output multi-dimensional diagnostic intermediate results.
[0106] 3) Build a localized intelligent decision-making closed loop: Deploy lightweight deep models (such as DeepSeek) locally, generate executable operation and maintenance suggestions (such as "prioritize expanding the memory of node A today") based on historical data and real-time analysis results, and automatically generate natural language reports that conform to enterprise standards, reducing reliance on human experience.
[0107] 4) Ensure data security throughout the entire process: All data processing and model inference are completed locally to prevent sensitive logs from being leaked to the cloud. At the same time, access control and encrypted storage mechanisms ensure data compliance in the intranet environment, balancing functional requirements and security requirements.
[0108] In one specific embodiment, the automatic inspection method is applied to an automatic inspection system, the architecture of which is shown in the figure below. Figure 2 As shown, it includes four core modules: automatic inspection execution, log collection and analysis, local intelligent processing, and report generation and output. The descriptions of each module are as follows:
[0109] (1) Automatic inspection execution module: Control multiple devices to log in concurrently through SecureCRT scripts, execute inspection commands (such as show cpu, display memory), and store the raw logs in the local database to avoid delays caused by manual operation.
[0110] (2) Log collection and analysis module: Standardize and clean the raw logs (such as removing garbled characters and aligning timestamps), extract key indicators (such as regular expression matching CPU utilization: (\d+)%), and locate anomalies through rule engine (such as threshold alarm) or statistical model (such as time series anomaly detection).
[0111] (3) Local intelligent processing module: Input structured data into a locally deployed fine-tuned NLP (Natural Language Processing) model (such as the DeepSeek model), and combine it with time series data in the historical database to generate operation and maintenance suggestions (such as "Node A's memory usage has exceeded 80% for 3 consecutive days, and it is recommended to expand its capacity").
[0112] (4) Report generation and output module: Automatically generates reports containing charts and natural language descriptions (e.g., Markdown to PDF), notifies maintenance personnel via email or message, and updates the local knowledge base. Report templates are configurable and support multilingual generation; the notification module integrates enterprise IM (Instant Messaging) interfaces (e.g., DingTalk / WeChat Work).
[0113] Based on this automatic inspection system, the corresponding automatic inspection method includes the following steps:
[0114] S1, Automatic Inspection Execution
[0115] 1. SecureCRT scripts control multi-device login
[0116] Script Development: Use the Python scripting language supported by SecureCRT to write automated login scripts that connect to target devices (such as switches and servers) via SSH / Telnet protocols.
[0117] Device management: Store device IP, port, account password in a local encrypted configuration file. The script dynamically reads and adapts to the authentication methods of different devices (such as RSA key, two-factor authentication).
[0118] Multi-threaded control: Concurrent login to multiple devices is achieved through thread pool technology, and a timeout retry mechanism is set (such as 3 retries with a 10-second interval) to avoid single point of failure causing the overall process to be interrupted.
[0119] 2. Execute inspection commands concurrently
[0120] Command library configuration: Pre-set commonly used inspection commands (such as show cpu usage, display interfacebrief), and automatically switch syntax according to device type (Huawei / Cisco).
[0121] Command distribution: After successful login, the script sends commands to the device in a preset order and matches the command response status (such as Success / Error) in real time using regular expressions.
[0122] Error handling: If a command execution failure is detected (such as connection interruption or syntax error), the error log will be automatically recorded and a retry will be triggered or the device will be marked as "requiring manual intervention".
[0123] 3. Collect logs to the local database
[0124] Log capture: The script captures the raw log output of the SecureCRT terminal in real time, removes control characters (such as ANSI color codes), and retains plain text data.
[0125] Structured storage: Logs are categorized by device IP, timestamp, and command type and stored in a local relational database (such as MySQL) or time-series database (such as InfluxDB), with indexes created to accelerate subsequent queries.
[0126] Data encryption: Sensitive information (such as device passwords and log content) is stored using AES-256 (Advanced Encryption Standard-256bit), and the key is managed through HSM (Hardware Security Module).
[0127] It should be noted that the entire log content can be encrypted, or only some sensitive fields in the log can be encrypted.
[0128] S2, Log Collection and Analysis
[0129] 1. Log preprocessing (normalization / denoising)
[0130] Format cleaning: Use regular expressions to remove irrelevant characters (such as garbled characters and repeated spaces) from the logs and extract structured fields (such as CPU utilization: 75%).
[0131] Timestamp alignment: Convert timestamps from different devices to UTC format to resolve time zone inconsistencies.
[0132] Noise filtering: Based on a preset rule base (such as whitelist keywords), discard irrelevant log lines (such as debugging information).
[0133] 2. Key Indicator Extraction (Regular Expression)
[0134] Metric template configuration: Define metric extraction rules for each type of device (e.g., Huawei device CPU utilization regular expression: CPUusage.*?(\d+)%).
[0135] Dynamic parsing: Call Python's re library or open-source parsing tools (such as Grok) to match key metrics in the logs (such as memory usage and port traffic).
[0136] Data normalization: The extracted values are standardized into standardized units (such as MB, Gbps) to facilitate subsequent analysis.
[0137] It should be noted that the key metrics are quantifiable values (such as CPU utilization, port traffic, etc.).
[0138] 3. Anomaly Detection and Statistics
[0139] Threshold alarm: Based on a preset rule engine (such as Drools), trigger a static threshold alarm (such as CPU>90% for 5 minutes).
[0140] Dynamic baseline detection: Calculates historical data baselines using statistical models (such as Z-score, moving average) to identify abnormal fluctuations that deviate from the normal range.
[0141] Association analysis: Use graph databases (such as Neo4j) to build device topology relationships and detect cross-device association anomalies (such as packet loss on a switch port causing increased latency on downstream servers).
[0142] It should be noted that anomaly detection and statistics employ a dual detection mechanism of "static threshold + dynamic baseline". Dynamic baseline detection calculates historical data baselines using a statistical model to identify latent trend anomalies; static threshold alarms, based on a preset rule engine, are used to capture sudden exceedances. Furthermore, correlation analysis is used, employing a graph database to construct device topology relationships, performing correlation analysis on anomalies of individual devices to discover causal relationships across devices.
[0143] S3, Local Intelligent Processing
[0144] 1. Input the local DeepSeek processing model
[0145] Model deployment: Package the fine-tuned NLP model (such as DeepSeek based on BERT) into a Docker service and deploy it on a local server or edge device.
[0146] Data input: Concatenate structured metrics (JSON format) with time-series data from the historical database to form the model input vector.
[0147] Model Inference: Call the local inference interface to output the root cause diagnosis results and confidence scores.
[0148] 2. Analyze correlations with historical data
[0149] Time-series data retrieval: Query the index trends of the same device over the past 30 days from the database to identify periodic patterns (such as CPU load during daily peak hours).
[0150] Root cause reasoning: Using causal inference algorithms (such as the PC algorithm) to analyze the dependencies between abnormal indicators and locate the root cause (such as a configuration change causing port congestion).
[0151] Knowledge graph query: Based on a pre-built operation and maintenance knowledge graph, match solutions for similar historical faults.
[0152] It should be noted that the operation and maintenance knowledge graph integrates multi-dimensional information such as equipment type, fault type, and indicator parameters. It not only stores a massive amount of historical fault cases and solutions, but also quickly matches similar problems through graph association, providing operation and maintenance personnel with efficient handling suggestions and realizing the structured reuse of experience.
[0153] 3. Generate operation and maintenance suggestions
[0154] Strategy template library: Pre-set repair strategies for common scenarios (such as "expand memory" and "reboot port"), and match the best solution based on the model output.
[0155] Natural Language Generation: Use template engines (such as Jinja2) or lightweight NLP models (such as T5) to convert structured recommendations (i.e., remediation strategies) into natural language descriptions (such as "Recommend restarting port GigabitEthernet0 / 1 after 17:00 today").
[0156] Compliance verification: The rules engine checks whether the recommendations comply with the company's operational guidelines (such as prohibiting operations during peak business hours).
[0157] S4. Report Generation and Output
[0158] 1. Generate a structured analysis report
[0159] Template engine call: Use Word templates to insert dynamic data (such as tables and line charts).
[0160] Visualization rendering: Generate charts (such as CPU trend charts and anomaly distribution pie charts) using Matplotlib or Echarts and embed them in reports.
[0161] Multiple export formats: Supports PDF, HTML, and Excel formats to suit different reading scenarios.
[0162] 2. Output diagnostic results
[0163] Prioritization: Categorize diagnostic results by fault level (urgent / important / warning), highlighting high-risk items.
[0164] Evidence chain correlation: Mark the source of abnormal data in the report (e.g., "based on the log of device 192.168.1.1 at 14:00 on 2023-10-01") to ensure that the conclusions are traceable.
[0165] 3. Notify maintenance personnel
[0166] Push notifications: Integrate with enterprise IM robots (such as DingTalk / Lark) to send report links and summaries to designated groups via Webhook.
[0167] Email notification: Sends an encrypted email using the SMTP protocol, with the full report and execution suggestion ticket attached.
[0168] Alarm escalation: If an emergency fault is detected (such as a core device failure), an automatic telephone alarm (such as TwilioAPI) will be triggered.
[0169] It should be noted that the automatic inspection method provided by this invention has the following characteristics:
[0170] a) Improved efficiency and accuracy:
[0171] Multi-device concurrent login and command control based on SecureCRT scripts transforms manual operation of each device into automated execution, reducing inspection time by more than 80% and avoiding the risk of human error.
[0172] By preprocessing logs (standardization / denoising) and extracting key metrics (regular expression matching), the raw logs are transformed into structured data, effectively filtering usable information and providing high-precision input for subsequent analysis.
[0173] b) Intelligent operation and maintenance decision support:
[0174] By combining localized DeepSeek model inference with historical data correlation analysis, the system can locate the root cause of anomalies and generate operation and maintenance suggestions (such as scaling strategies and configuration optimization), reducing reliance on human experience and improving diagnostic accuracy by 60%.
[0175] The dynamic anomaly detection mechanism (threshold alarm + statistical model) can identify potential faults (such as memory leaks and port congestion) in real time, support proactive early warning, and avoid business interruption.
[0176] c) Security and compliance assurance:
[0177] The entire data process is handled in a closed loop within the intranet environment, avoiding the risk of sensitive information leakage caused by cloud transmission and meeting the strong compliance requirements of industries such as finance and government.
[0178] The report generation module ensures rapid response from operations and maintenance personnel through structured analysis reports and automatic notifications (email / IM), while audit logs are stored in encrypted form for easy accountability.
[0179] The automated inspection method provided in this invention first automatically logs into multiple devices using a pre-written automated login script in terminal simulation software and executes inspection commands. Then, it acquires the raw logs output by the multiple devices after executing the inspection commands and stores these logs in a local database. Next, it preprocesses the stored raw logs and extracts key indicators from them using preset indicator extraction rules. It then performs anomaly detection on the extracted key indicators, obtaining abnormal indicators and cross-device correlation anomaly results for the corresponding devices. Based on the abnormal indicators and the cross-device correlation anomaly results, it outputs root cause diagnosis results using a fine-tuned Natural Language Processing (NLP) model combined with the historical time-series data of the corresponding devices. The fine-tuned NLP model is deployed on a local server. Finally, it obtains corresponding maintenance suggestions based on the root cause diagnosis results and a pre-built maintenance knowledge graph, and outputs a structured analysis report. This invention achieves automated inspection of multiple devices through an automated login script, improving inspection efficiency and reducing manual operations. Meanwhile, by extracting key indicators and utilizing anomaly detection technology and a fine-tuned NLP model combined with historical data to output root cause diagnosis results, the intelligence level of log analysis is significantly improved. Furthermore, generating operational suggestions based on an operational knowledge graph enhances intelligent assistance functions. By storing raw logs in a local database and deploying the fine-tuned NLP model on a local server, end-to-end data security is ensured, achieving automated and intelligent operational management. This addresses the problems of existing inspection technologies, which primarily rely on manual operation or basic script tools, resulting in low efficiency of manual inspections, insufficient intelligence in log analysis, lack of intelligent assistance, and security risks in data flow.
[0180] Example 2:
[0181] like Figure 3 As shown, this embodiment provides an automatic inspection device for performing the above-described automatic inspection method, including:
[0182] The automatic login inspection module 11 is used to automatically log in to multiple devices and execute inspection commands through a pre-written automated login script of terminal emulation software.
[0183] The raw log storage module 12 is connected to the automatic login inspection module 11 and is used to obtain the raw logs output by multiple devices after executing the inspection command, and store the raw logs in the local database.
[0184] The key indicator extraction module 13 is connected to the original log storage module 12 and is used to preprocess the stored original log and extract key indicators from the preprocessed original log using preset indicator extraction rules.
[0185] Anomaly detection module 14, connected to key indicator extraction module 13, is used to perform anomaly detection on the extracted key indicators and obtain the abnormal indicators and the cross-device correlation anomaly results of the corresponding devices.
[0186] The root cause diagnosis module 15 is connected to the anomaly detection module 14 and is used to output the root cause diagnosis result based on the anomaly indicators and the cross-device correlation anomaly results of the corresponding devices, using a fine-tuned natural language processing (NLP) model combined with the historical time series data of the corresponding devices. The fine-tuned NLP model is deployed on a local server.
[0187] The operation and maintenance suggestion module 16 is connected to the root cause diagnosis module 15 and is used to obtain corresponding operation and maintenance suggestions based on the root cause diagnosis results and the pre-built operation and maintenance knowledge graph, and output a structured analysis report.
[0188] Optionally, the terminal emulation software is SecureCRT, and the automatic login inspection module 11 includes:
[0189] The concurrent login unit is used to concurrently log in to multiple devices using thread pool technology, through a pre-written SecureCRT automated login script and pre-encrypted and stored account and password for each device.
[0190] The inspection command sending unit is used to send inspection commands to multiple devices in a preset order.
[0191] Optionally, the raw log storage module 12 includes:
[0192] The raw log acquisition unit is used to capture the raw logs output by multiple devices after executing the inspection command in real time, and remove control characters while retaining plain text data.
[0193] The categorized storage unit is used to categorize and store raw log data of plain text into a local relational database or time-series database.
[0194] Optionally, the key indicator extraction module 13 includes:
[0195] The format cleaning unit is used to clean the format of the stored raw log to remove garbled characters and duplicate spaces in the raw log.
[0196] A unified timestamp unit is used to convert the timestamps in the original logs of multiple devices after format cleaning into Coordinated Universal Time (UTC) format.
[0197] Optionally, the anomaly detection module 14 includes:
[0198] A normalization unit is used to normalize the extracted key indicators;
[0199] The first abnormal indicator acquisition unit is used to perform threshold evaluation on the normalized key indicators based on the preset rule engine to obtain the first abnormal indicator that triggers the static threshold alarm.
[0200] The second abnormal indicator acquisition unit is used to calculate the historical data baseline through a preset statistical model and identify the second abnormal indicator that deviates from the normal range among the key indicators after normalization.
[0201] The correlation anomaly detection unit is used to detect cross-device correlation anomalies based on the devices corresponding to the first and second anomaly indicators, using a pre-built graph database-based device topology relationship, and obtain the corresponding results.
[0202] Optionally, the device further includes:
[0203] The model encapsulation and deployment module is used to encapsulate the fine-tuned NLP model as a Docker service and deploy it on a local server.
[0204] Optionally, the operation and maintenance suggestion module 16 includes:
[0205] The solution matching unit is used to match historical faults similar to the root cause diagnosis results and corresponding solutions based on a pre-built operation and maintenance knowledge graph.
[0206] The operation and maintenance suggestion acquisition unit is used to determine the corresponding repair strategy based on the solution, and convert the repair strategy into a natural language description to obtain the corresponding operation and maintenance suggestion.
[0207] The report generation unit is used to generate a corresponding structured analysis report based on the operation and maintenance recommendations using a template engine.
[0208] Example 3:
[0209] refer to Figure 4This embodiment provides an automatic inspection device, including a memory 21 and a processor 22. The memory 21 stores a computer program, and the processor 22 is configured to run the computer program to execute the automatic inspection method in Embodiment 1.
[0210] The memory 21 is connected to the processor 22. The memory 21 can be a flash memory, a read-only memory or other memory, and the processor 22 can be a central processing unit or a microcontroller.
[0211] Example 4:
[0212] This embodiment provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the automatic inspection method in Embodiment 1 above.
[0213] The computer-readable storage medium includes volatile or non-volatile, removable or non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, computer program modules, or other data). Computer-readable storage media include, but are not limited to, RAM (Random Access Memory), ROM (Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), flash memory or other memory technologies, CD-ROM (Compact Disc Read-Only Memory), DVD or other optical disc storage, cartridges, magnetic tapes, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer.
[0214] In summary, the automatic inspection method, apparatus, and readable storage medium provided in this invention first automatically logs into multiple devices using a pre-written automated login script of terminal simulation software and executes inspection commands. Then, it acquires the raw logs output by the multiple devices after executing the inspection commands and stores these raw logs in a local database. Next, it preprocesses the stored raw logs and extracts key indicators from the preprocessed logs using preset indicator extraction rules. It then performs anomaly detection on the extracted key indicators, obtaining abnormal indicators and cross-device correlation anomaly results for the corresponding devices. Based on the abnormal indicators and the cross-device correlation anomaly results, it outputs root cause diagnosis results using a fine-tuned Natural Language Processing (NLP) model combined with the historical time-series data of the corresponding devices. The fine-tuned NLP model is deployed on a local server. Finally, it obtains corresponding maintenance suggestions based on the root cause diagnosis results and a pre-built maintenance knowledge graph, and outputs a structured analysis report. This invention achieves automatic inspection of multiple devices through an automated login script, improving inspection efficiency and reducing manual operations. Meanwhile, by extracting key indicators and utilizing anomaly detection technology and a fine-tuned NLP model combined with historical data to output root cause diagnosis results, the intelligence level of log analysis is significantly improved. Furthermore, generating operational suggestions based on an operational knowledge graph enhances intelligent assistance functions. By storing raw logs in a local database and deploying the fine-tuned NLP model on a local server, end-to-end data security is ensured, achieving automated and intelligent operational management. This addresses the problems of existing inspection technologies, which primarily rely on manual operation or basic script tools, resulting in low efficiency of manual inspections, insufficient intelligence in log analysis, lack of intelligent assistance, and security risks in data flow.
[0215] It is understood that the above embodiments are merely exemplary implementations used to illustrate the principles of the present invention, and the present invention is not limited thereto. For those skilled in the art, various modifications and improvements can be made without departing from the spirit and essence of the present invention, and these modifications and improvements are also considered to be within the scope of protection of the present invention.
Claims
1. An automatic inspection method, characterized in that, The method includes: The system automatically logs into multiple devices and executes inspection commands using a pre-written automated login script for terminal emulation software. Obtain the raw logs output by multiple devices after executing the inspection command, and store the raw logs in a local database; The stored raw logs are preprocessed, and key indicators are extracted from the preprocessed raw logs using preset indicator extraction rules. Anomaly detection is performed on the extracted key indicators to obtain the abnormal indicators and the cross-device correlation anomaly results of the corresponding devices; Based on the aforementioned abnormal indicators and the cross-device correlation abnormal results of the corresponding devices, the root cause diagnosis results are output using a fine-tuned natural language processing (NLP) model combined with the historical time series data of the corresponding devices. The fine-tuned NLP model is deployed on a local server. Based on the root cause diagnosis results and the pre-built operation and maintenance knowledge graph, corresponding operation and maintenance suggestions are obtained, and a structured analysis report is output.
2. The method according to claim 1, characterized in that, The terminal emulation software is SecureCRT. The automated login script of the pre-written terminal emulation software automatically logs into multiple devices and executes inspection commands, specifically including: The system uses a thread pool to concurrently log in to multiple devices using a pre-written SecureCRT automated login script and pre-encrypted and stored usernames and passwords for each device. Inspection commands are sent to multiple devices in a preset order.
3. The method according to claim 1, characterized in that, The step of obtaining the raw logs output by multiple devices after executing the inspection command and storing the raw logs in a local database specifically includes: The system captures the raw logs output by multiple devices after executing the inspection command in real time, removes control characters, and retains plain text data. The raw log data of plain text is categorized and stored in a local relational database or time-series database.
4. The method according to claim 1, characterized in that, The preprocessing of the stored raw logs specifically includes: The stored raw logs are formatted and cleaned to remove garbled characters and duplicate spaces. The timestamps in the original logs from multiple devices, after being format-cleaned, are uniformly converted to Coordinated Universal Time (UTC) format.
5. The method according to claim 1, characterized in that, The step of performing anomaly detection on the extracted key indicators to obtain abnormal indicators and corresponding cross-device correlation anomaly results for the devices specifically includes: The extracted key indicators are then normalized. Based on the preset rule engine, the threshold evaluation of the normalized key indicators is performed to obtain the first abnormal indicator that triggers the static threshold alarm. By calculating the historical data baseline using a pre-set statistical model, the second abnormal indicator that deviates from the normal range among the key indicators after normalization is identified. Based on the devices corresponding to the first and second anomaly indicators, cross-device association anomalies are detected using a pre-built graph database-based device topology relationship, and corresponding results are obtained.
6. The method according to claim 1, characterized in that, Before outputting the root cause diagnosis result based on the abnormal indicators and the cross-device correlation anomaly results of the corresponding devices using a fine-tuned Natural Language Processing (NLP) model combined with the historical time-series data of the corresponding devices, the method further includes: The fine-tuned NLP model is packaged into a Docker service and deployed on a local server.
7. The method according to claim 6, characterized in that, The process involves obtaining corresponding operational suggestions based on the root cause diagnosis results and a pre-constructed operational knowledge graph, and outputting a structured analysis report, specifically including: Based on a pre-built operation and maintenance knowledge graph, historical faults similar to the root cause diagnosis results and corresponding solutions are matched. Based on the solution, a corresponding repair strategy is determined, and the repair strategy is converted into a natural language description to obtain corresponding operation and maintenance suggestions; Based on the aforementioned maintenance recommendations, a corresponding structured analysis report is generated using a template engine.
8. An automatic inspection device, characterized in that, The device includes: The automatic login inspection module is used to automatically log in to multiple devices and execute inspection commands through a pre-written automated login script of terminal emulation software. The raw log storage module is connected to the automatic login inspection module and is used to obtain the raw logs output by multiple devices after executing the inspection command, and store the raw logs in a local database. The key indicator extraction module is connected to the original log storage module and is used to preprocess the stored original logs and extract key indicators from the preprocessed original logs using preset indicator extraction rules. An anomaly detection module, connected to the key indicator extraction module, is used to perform anomaly detection on the extracted key indicators and obtain the anomaly indicators and the cross-device associated anomaly results of the corresponding devices. The root cause diagnosis module, connected to the anomaly detection module, is used to output the root cause diagnosis result based on the anomaly indicators and the cross-device correlation anomaly results of the corresponding devices, using a fine-tuned natural language processing (NLP) model combined with the historical time series data of the corresponding devices. The fine-tuned NLP model is deployed on a local server. The operation and maintenance suggestion module is connected to the root cause diagnosis module. It is used to obtain corresponding operation and maintenance suggestions based on the root cause diagnosis results and the pre-built operation and maintenance knowledge graph, and output a structured analysis report.
9. An automatic inspection device, characterized in that, It includes a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to implement the automatic inspection method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the automatic inspection method as described in any one of claims 1-7.
Citation Information
Cited By
Inspection data processing method and system
CN121349756A
Cross-platform RPA data inspection robot control system based on federal architecture
CN121907893A
A cross-platform RPA data inspection robot control system based on a federated architecture
CN121907893B
Root cause positioning method and device for memory overflow of database and medium
CN121979719A
Software exception handling method, exception handling system and storage medium
CN121996466A