Method, device, equipment, storage medium and program product for processing operation and maintenance command

By transforming operation and maintenance commands into operation behavior graph nodes and utilizing a graph neural network risk prediction model, the problem of high error rate in traditional operation and maintenance is solved, and accurate risk assessment and security control are achieved in a multi-platform, multi-role, and multi-task parallel environment.

CN120872679BActive Publication Date: 2026-01-27INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511403625.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-28
Publication Date
2026-01-27
Estimated Expiration
2045-09-28

AI Technical Summary

Technical Problem

Traditional operations and maintenance rely on manual command execution, which leads to problems such as configuration errors, missing parameters, and improper command order. It is difficult to adapt to dynamic operations and maintenance environments with multiple platforms, multiple roles, and multiple tasks running in parallel, and it cannot provide accurate and real-time operational risk assessments and intervention suggestions.

Method used

Operation and maintenance commands are transformed into nodes in an operational behavior graph. The graph records the relationships between metadata, and a risk prediction model using a graph neural network is used to calculate risk probabilities. Based on the risk probabilities, security control strategies are determined, and real-time risk assessment and intervention suggestions are provided.

Benefits of technology

It reduces reliance on human experience, avoids configuration errors and improper command sequences, and enables accurate risk assessment and security control in dynamic environments, adapting to multi-platform, multi-role, and multi-task parallel operation and maintenance scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120872679B_ABST
    Figure CN120872679B_ABST
Patent Text Reader

Abstract

The application discloses an operation and maintenance command processing method and device, equipment, a storage medium and a program product, relates to the technical field of server operation and maintenance, and comprises the following steps: converting an operation and maintenance command into a node in an operation and maintenance command operation behavior graph, recording the association relationship between metadata in the graph, reducing the dependence on manual experience, and avoiding configuration errors, improper command sequences and the like through structured graph analysis; a risk prediction model based on a graph neural network can utilize the information of operation and maintenance command metadata and the association relationship in the graph to perform accurate risk probability calculation on a command to be executed in a dynamic and complex environment, break through the limitations of traditional static rules; a safety control strategy is determined through the risk probability, real-time risk assessment and intervention suggestions can be provided before the command is executed, system interruption, data anomaly and the like can be effectively prevented, and the dynamic operation and maintenance environment of multiple platforms, multiple roles and multiple tasks in parallel is adapted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of server operation and maintenance, and in particular to a method, apparatus, device, storage medium, and program product for processing operation and maintenance commands. Background Technology

[0002] As server clusters continue to expand and automation and intelligence levels increase, enterprise data centers, financial cloud platforms, and core information systems face increasingly stringent requirements for system stability and availability. Routine maintenance has become a crucial element in ensuring continuous system operation. However, traditional maintenance still heavily relies on manual command execution, experience-based judgment, and script management, making it highly susceptible to errors in complex environments. Issues such as configuration errors, missing parameters, and improper command order not only frequently lead to system interruptions and data anomalies but may even trigger cascading failures or security incidents. Especially in multi-platform, multi-role, and multi-task parallel maintenance environments, traditional static rule-based maintenance auditing and approval mechanisms are ill-suited to highly dynamic operational behaviors and diverse risks, failing to provide accurate, real-time, and context-sensitive operational risk assessments and intervention recommendations. Summary of the Invention

[0003] This application provides a method, apparatus, device, storage medium, and program product for processing operation and maintenance commands, in order to at least solve the problem of high error rate in related operation and maintenance technologies.

[0004] This application provides a method for processing operation and maintenance commands, including: obtaining an operation and maintenance command to be executed; determining the target metadata corresponding to the operation and maintenance command to be executed based on a pre-constructed operation behavior graph; wherein, the operation behavior graph uses the metadata of the operation behavior events of the operation and maintenance command as nodes and the association relationship between the metadata as edges; obtaining the risk probability of the operation and maintenance command to be executed output by the risk prediction model based on the target metadata and the risk prediction model; wherein, the risk prediction model is constructed based on a graph neural network and trained from sample commands and their risk probabilities; and determining the security control strategy of the operation and maintenance command to be executed based on the risk probability.

[0005] This application also provides a processing device for operation and maintenance commands, including:

[0006] The acquisition module is used to acquire operation and maintenance commands to be executed.

[0007] The graph indexing module is used to determine the target metadata corresponding to the operation and maintenance command to be executed based on the pre-constructed operation behavior graph; wherein, the operation behavior graph uses the metadata of the operation behavior events of the operation and maintenance command as nodes and the association relationship between the metadata as edges;

[0008] The risk probability prediction module is used to obtain the risk probability of the operation and maintenance command to be executed, output by the risk prediction model, based on the target metadata and the risk prediction model; wherein, the risk prediction model is constructed based on a graph neural network and is trained from sample commands and their risk probabilities;

[0009] The control strategy determination module is used to determine the security control strategy for the operation and maintenance command to be executed based on the risk probability.

[0010] This application also provides an electronic device, including: a memory for storing a computer program; and a processor for executing the steps of the above-described operation and maintenance command processing method when executing the computer program.

[0011] This application also provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, it implements the steps of the above-mentioned operation and maintenance command processing method.

[0012] This application also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the above-mentioned operation and maintenance command processing method.

[0013] This application transforms operation and maintenance commands into nodes in an operational behavior graph. By recording the relationships between metadata in the graph, it reduces reliance on human experience. At the same time, structured graph analysis avoids problems such as configuration errors and improper command order. Based on a graph neural network-based risk prediction model, it can use the information of nodes (operation and maintenance command metadata) and edges (relationships) in the graph to accurately calculate the risk probability of commands to be executed in dynamic and complex environments, breaking through the limitations of traditional static rules. By determining security control strategies through risk probability, it can provide real-time risk assessment and intervention suggestions before command execution, effectively preventing problems such as system interruption and data anomalies, and adapting to dynamic operation and maintenance environments with multiple platforms, multiple roles, and multiple tasks running in parallel. Attached Figure Description

[0014] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0015] Figure 1 A schematic diagram of the specific hardware architecture upon which the execution of operation and maintenance commands depends;

[0016] Figure 2 A flowchart illustrating a method for processing operation and maintenance commands provided in an embodiment of this application;

[0017] Figure 3 A flowchart illustrating another method for processing operation and maintenance commands provided in this application embodiment;

[0018] Figure 4 A schematic diagram of the structure of a device for processing operation and maintenance commands provided in an embodiment of this application;

[0019] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0020] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0021] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0022] To more clearly illustrate the embodiments of this application, the technical terms used in the embodiments will be briefly introduced below:

[0023] Graph Neural Networks (GNNs) are a class of deep learning models specifically designed for processing graph-structured data.

[0024] A multilayer perceptron (MLP) is a feedforward artificial neural network that maps a set of input vectors to a set of output vectors.

[0025] Graph tensors represent and process graph-structured data in the form of tensors. A tensor is a generalized array that can be a zero-dimensional scalar, a one-dimensional vector, a two-dimensional matrix, or even a higher-dimensional array. In graph tensors, various information about the graph is encoded.

[0026] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0027] like Figure 1As shown, Figure 1 A schematic diagram of the specific hardware architecture upon which the execution of operation and maintenance commands depends.

[0028] The hardware architecture includes: an operational data modeling module, a risk assessment module, an auxiliary decision generation module, an execution control module, and a feedback learning and optimization module.

[0029] The operation data modeling module continuously collects operation behaviors and their contextual information, such as command content, execution path, user identity, and system load status, from sources such as command terminals, operation logs, and system monitoring. It constructs a structured operation behavior graph to depict the dependency relationship between commands and the environment, and realizes the joint expression of operation intentions and environment status.

[0030] The risk assessment module uses a graph neural network to embed and calculate the graph nodes for the current operation and maintenance command to be analyzed, outputs the risk probability score of the operation, and determines the risk level (high, medium, or low) based on the model output.

[0031] The decision support generation module automatically generates operation prompts and suggestions based on the evaluation results, including confirmation prompts, command alternative recommendations, parameter suggestions, and risk descriptions, to achieve real-time and actionable intelligent reminders and help operation and maintenance personnel make safer operation choices.

[0032] The execution control module triggers different levels of security control policies based on the risk level. For example, high-risk operations require dual approval, enable semantic interpretation, and force delayed confirmation; medium-risk operations provide rollback suggestions; and low-risk operations are directly allowed and recorded.

[0033] The feedback learning and optimization module tracks the actual consequences of executed commands and compares them with risk prediction results, automatically updating risk model parameters and prompt strategies to build a closed-loop self-learning system.

[0034] The embodiments of this application provide a method for processing operation and maintenance commands. The method is described in detail below in conjunction with the execution flow of the operation and maintenance command processing method.

[0035] like Figure 2 As shown, Figure 2 The flowchart illustrates a method for processing operation and maintenance commands provided in this application embodiment. The method includes the following steps:

[0036] S201. Obtain the operation and maintenance commands to be executed.

[0037] Operation and maintenance commands to be executed are instructions used to perform core operation and maintenance tasks during the lifecycle of servers and other infrastructure. These core tasks include deployment, monitoring, configuration, troubleshooting, and resource management. Examples of commands to be executed include "iptables -F" for configuring IPv4 packet filtering rules and "cat / etc / passwd" for viewing user information.

[0038] S202. Based on the pre-built operation behavior map, determine the target metadata corresponding to the operation and maintenance command to be executed.

[0039] The operation behavior graph uses the metadata of operation behavior events of maintenance commands as nodes and the relationships between metadata as edges. Metadata includes at least one of the following: command content, execution path, user identity, system load status, preceding command sequence, and host identifier. Relationships include, but are not limited to, preceding execution, dependency, and context. A pre-built operation behavior graph is used to depict the dependencies between commands and the environment.

[0040] The target metadata corresponding to the operation and maintenance command to be executed includes the metadata of the operation behavior event of the operation and maintenance command to be executed, as well as the metadata associated with that metadata.

[0041] In some embodiments, the process of constructing the operation behavior graph includes: obtaining metadata and context information of operation behavior events of operation and maintenance commands, then constructing structured operation behavior events based on the metadata, and then constructing the operation behavior graph based on the structured operation behavior events and context information.

[0042] Specifically, the operation and maintenance command-related data is first collected from multiple data sources. These data sources include, but are not limited to, command-line terminals, the current directory, the executing user, system load status, preceding commands, and target system roles. The operation and maintenance command-related data includes terminal command input (command), current execution path (path), system load status (system_load), operating user (user), and preceding command sequence (prev_commands[]), etc.

[0043] Optionally, the original operation and maintenance command-related data is cleaned, including: removing invalid commands, such as empty inputs and incorrectly formatted instructions; supplementing missing context information, such as matching isolated commands with corresponding users and system states; standardizing data formats, such as unifying path representation and converting system load values ​​into standardized numerical values; ultimately forming a structured operation behavior event dataset, providing high-quality input for subsequent graph construction. The structured operation behavior event dataset is shown in Table 1, including the data source, operation and maintenance command-related data, and example values.

[0044] Table 1

[0045]

[0046] Furthermore, an operational behavior graph G=(V,E) is constructed based on structured operational behavior events. Based on the cleaned data, the core elements involved in the operational behavior are abstracted into graph nodes V={v1,v2,……,v...} n} represents operation behavior events, such as command nodes that store specific command text and semantic information, user nodes that record the identity and permissions of the user performing the operation, path nodes that represent the file system path that represents the function of the command, and system status nodes that reflect the system operating parameters when the command is executed.

[0047] Based on the semantic relationships and operational logic between nodes, define and generate the edges E={e} of the graph. ij} represents the relationships between nodes, including preorder execution relationships, operation object relationships, execution subject relationships, and environment relationships. The initial weight of the edges can be set according to the strength of the association; for example, commands and paths that frequently appear together have higher weights. This can be dynamically adjusted through feedback learning. By constructing edge relationships, isolated nodes are connected into a logically related graph structure, forming a complete operational behavior graph G=(V,E).

[0048] The relationships between nodes are shown in Table 2.

[0049] Table 2

[0050]

[0051] As shown in Table 2, the preceding execution relationship connects the current command node (command_node) with the preceding command sequence node (prev_command_node), representing the temporal dependency of command execution; the operation object relationship connects the current command node with the path node (path_node), reflecting the effect of the command on the target path; the execution subject relationship connects the current command node with the user node (user_node), clarifying the initiator of the command; and the environment relationship connects the command node with the system state node (system_load_node), reflecting the system environment of command execution. The semantic vector of the command node, transformed by the pre-trained model, retains the dangerous attribute characteristics of the command, while the vectors of the user, path, and system state nodes quantify the risk weights of environmental factors.

[0052] Optionally, the constructed behavioral graph containing node vectors, edge relationships, and weights can be stored in a graph database to achieve persistent storage and efficient querying of the graph. Simultaneously, to meet the inference requirements of the graph neural network, the graph data in the graph database is periodically converted into graph tensor format. This involves arranging node vectors sequentially to form a node matrix, representing edge relationships and weights as adjacency matrices, and finally generating numerical tensor data that can be directly input into the GNN model, providing structured input for subsequent node embedding calculations and risk scoring in the risk assessment module.

[0053] The above embodiments clean and integrate heterogeneous data such as terminal commands, execution paths, user permissions, and system load, abstract scattered operational elements into nodes and define their relationships to construct a structured graph. This not only retains the attribute characteristics of individual elements but also captures the dependency logic between elements through edge relationships, providing structured input for risk reasoning in graph neural networks. This enables risk assessment to no longer rely on single command features but can output accurate risk probabilities by combining complete context.

[0054] Once the operation behavior graph is constructed, when an operation and maintenance command to be executed is obtained, the target metadata corresponding to the operation and maintenance command to be executed is determined based on the operation behavior graph.

[0055] S203. Based on the target metadata and the risk prediction model, obtain the risk probability of the operation and maintenance command to be executed output by the risk prediction model.

[0056] The risk prediction model is based on a graph neural network and is trained from sample commands and their risk probabilities.

[0057] Optionally, the target metadata corresponding to the operation and maintenance command to be executed is vectorized, and the node (i.e., target metadata) v i Represented as a vector This serves as the input to the risk prediction model. For the first The initial embedding vectors of each node are the starting point of the model input. for A real space of dimensional numbers, that is, this vector has Composed of floating-point numbers, such as .

[0058] If node v iIf the command is the input, then a text vector is extracted from the command and used as input to the risk prediction model. A pre-trained word-to-vector (word2vec) model or a Transformer-based Bidirectional Encoder Representations from Transformers (BERT) model can be used to extract the text vector and map it to a fixed dimension. Filter out irrelevant features and enhance risk-related semantics.

[0059] word2vec learns the co-occurrence relationships of words in commands through a context window. For example, "rm" often appears together with "-rf" and "directory path," capturing the command's operation pattern. BERT, on the other hand, understands the complete semantics of commands through a bidirectional Transformer structure, such as distinguishing the difference in danger between "rmfile.txt" and "rm-rf / ".

[0060] The pre-trained word2vec models are shown in Table 3:

[0061] Table 3

[0062]

[0063] For example, the command `rm -rf / etc` can be converted to... The command `cat / etc / passwd` can be converted to... .

[0064] By using pre-trained models such as word2vec or BERT, keywords in commands (e.g., "rm", "-rf", "iptables"), operational logic (e.g., "delete", "clear rules"), and target objects (e.g., " / etc") are encoded into continuous vector space representations. This transforms the semantic differences between "force delete core directories" and "view file contents" into distance differences in vector space, providing a numerical foundation for subsequent models to capture command risk features. Fixed-dimensional vector representations ensure efficient inference and scalability of the risk prediction model.

[0065] If node If the user is an individual, then the user's permission level is determined, and this permission level is used as input to the risk prediction model. One-hot encoding or permission level mapping can be used; for example, root's permission level is 10, admin's permission level is 7, and guest's permission level is 2. For example, Root can be converted to [1.0,...,0.0], admin can be converted to [0.7,...,0.0], and guest can be converted to [0.2,...,0.0].

[0066] By using one-hot encoding or permission level mapping, abstract user identities and permissions are transformed into numerical features that can be directly processed by risk prediction models, providing a quantitative basis for subsequently capturing the risk correlation between users and commands. Intuitive permission level mapping and vector representation provide an interpretable foundation for risk assessment.

[0067] If node For the execution path, the execution path is divided into multiple sub-units. An embedding operation is performed on each sub-unit to obtain the embedding vector of each sub-unit. Then, the embedding vectors of each sub-unit are averaged to obtain the path-level feature vector of the execution path, which is used as the input of the risk prediction model. For example, / etc / network is converted to ['etc', 'network'], and after embedding, they are averaged as shown in the following formula (1):

[0068] (1)

[0069] For example, Commands can be converted to As the core directory of the system, “etc” has high-risk attributes in its embedding vector, while “network” represents the network configuration subdirectory, with the next highest risk attributes. The vector formed by averaging the two still highlights the core position of “etc”.

[0070] By breaking down execution paths into sub-units and embedding them separately, the risk attributes of key directory names can be captured individually. These sub-units are then integrated into a hierarchical path vector through averaging, preserving the semantic features of each level while achieving a unified representation of paths of different lengths. This provides fine-grained path features for subsequent risk assessment. The combined processing of hierarchical embedding and averaging aggregation ensures that the numerical distribution of the path vectors is highly correlated with the actual risk level, thereby enhancing the context-awareness of the overall risk assessment.

[0071] If node To determine the system load status (system_load), obtain the corresponding core resource and network metrics, including at least one of the following: processor load, memory usage, and network status, such as cpu_load=1.2, mem_usage=87%, net_tx=4.3MB / s. Then, standardize the numerical features and concatenate them into a system load status vector. Standardized processing using Z-scores or MinMax normalization eliminates dimensional differences between different metrics, enabling the effective fusion of multi-dimensional resource data. For example, Convert to [0.15, 0.86, ..., 0.38] to be used as input for the risk prediction model.

[0072] By extracting core indicators such as processor load, memory usage, and network status, the abstract system load status is transformed into quantifiable numerical characteristics, providing a quantitative basis for environmental factors in subsequent risk assessment.

[0073] If node This is a list of historical commands (prev_commands), which can be referenced from the command node and embedded by averaging historical data. For example, Convert to [0.23, 0.77, ..., 0.8].

[0074] The above embodiments employ differentiated vector initialization strategies for different types of nodes: command nodes extract text vectors using pre-trained word2vec or BERT models, which are then mapped to fixed-dimensional embedding vectors via a multilayer perceptron; user nodes are converted into low-dimensional vectors based on permission levels; path nodes use a hierarchical embedding method, splitting the path according to directory levels, embedding each level separately, and then averaging the results; system load status nodes standardize metrics such as Central Processing Unit (CPU) and memory, concatenating them into numerical vectors. Through these embodiments, unstructured element information is transformed into numerical node vectors that can be processed by graph neural networks (GNNs).

[0075] In some embodiments, the risk prediction model includes an embedding layer and a multilayer perceptron. When obtaining the risk probability corresponding to the command to be executed output by the risk prediction model based on the target metadata and the risk prediction model, the embedding layer first performs node-level embedding calculation on the target metadata to obtain the embedding vector of the target metadata; then, the multilayer perceptron performs probability calculation on the embedding vector to obtain the risk probability of the operation and maintenance command to be executed.

[0076] It can be understood that the target metadata first undergoes node-level embedding computation through the embedding layer. This is achieved by using a graph neural network to perform contextual association and information aggregation on the initial vectors of each node, generating an embedding vector that integrates multi-dimensional related features. Subsequently, a multilayer perceptron acts as a classifier, performing nonlinear transformations and mappings on the embedding vectors to output the risk probability of the operation and maintenance command to be executed.

[0077] Since the target metadata includes multiple types of nodes such as commands, users, paths, and system status, the embedding layer uses the GNN mechanism to allow node vectors to transmit information in the graph structure. This enables the final embedding vector to integrate multi-dimensional correlation information such as command semantics, user permissions, path importance, and system load status, thereby improving the completeness of risk features. The multilayer perceptron maps the embedding vector to probability values ​​in the [0,1] interval through nonlinear transformation, distinguishing various risk differences and capturing the risk fluctuations of the same command in different contexts, providing a precise basis for subsequent hierarchical control strategies.

[0078] The modular architecture of the aforementioned risk prediction model, combining an embedding layer and an MLP, balances feature representation capabilities with model efficiency, supporting dynamic system iteration. The embedding layer focuses on feature extraction and association modeling, and can be adjusted according to scenario requirements; the MLP, as a classifier, has a simple structure and efficient training, facilitating rapid parameter fine-tuning based on feedback data. This decoupled design allows the system to deeply mine risk features through the embedding layer and efficiently output risk probabilities through the MLP. Furthermore, it supports adding new metadata types by only extending the embedding layer processing logic, without reconstructing the overall model, thus improving system scalability and iteration efficiency.

[0079] Optionally, node-level embedding calculations are performed on the target metadata corresponding to the operation and maintenance command to be analyzed through the embedding layer, including: firstly, obtaining the neighbor node vectors of the target metadata in the adjacent upper layer through the embedding layer; for each neighbor node vector, performing weight transformation on the neighbor node vector according to the weight matrix of the adjacent upper layer, normalizing the weight-transformed neighbor node vectors and adding them together to obtain the sum; and then obtaining the embedding vector of the target metadata based on the sum and the activation function.

[0080] As shown in formula (2) below:

[0081] (2)

[0082] In formula (2), For the (l+1)th layer Let N be a set of node vectors, N(i) be the set of neighbors of the i-th node, and W be a set of node vectors. (l) Let W be the trainable weight matrix of the l-th layer of the graph neural network. (l) It is the core of the training model, which can automatically learn the risk contribution weights of different neighboring nodes to the current node during the training process; σ is the activation function, such as the Rectified Linear Unit (ReLU) or the Gaussian Error Linear Unit (GELU); d i d j Let be the degree of the node. Based on the node degree (d... i dj The normalization process avoids the problem of excessively large vector values ​​caused by too many neighbors, ensures that the features of different nodes have a balanced magnitude when aggregated, reduces the risk of gradient explosion during model training, and improves convergence efficiency.

[0083] This formula indicates that each node updates its own representation at each layer by aggregating information from its neighbors, thus achieving the layer-by-layer transmission of information in the graph. This process is equivalent to a command node absorbing information from its surroundings, causing the command's embedding vector to continuously integrate broader contextual relationships as the number of layers increases. After L layers of convolution, the embedding vector of the operation and maintenance command to be analyzed is obtained. This represents the overall semantic and potential risk characteristics of the operation in the current context.

[0084] Then, the embedding vector of the operation and maintenance commands to be analyzed Input the multilayer classifier (MLP) according to the following formula (3):

[0085] (3)

[0086] W in formula (3) clf Let be the classifier weight matrix, b be the bias term, and p∈[0,1] represent the probability that the current command is a dangerous operation. By using the sigmoid function to map this vector to the risk probability p in the interval [0,1], we can achieve both refined quantification of risk and provide a clear threshold basis for subsequent hierarchical control strategies.

[0087] The training process of the above-mentioned risk probability prediction model includes: the stage of training data collection and preprocessing, the stage of network structure initialization and parameter configuration, the stage of iterative training and parameter optimization, and the stage of continuous iterative training.

[0088] During the data collection and preprocessing phase, historical operation and maintenance data were collected in batches from multiple data sources, including command terminals, operation logs, and system monitoring. This included command content, execution path, user identity, system load status, and preceding command sequences. The raw data was then cleaned to form a structured operation behavior event dataset. Subsequently, an operation behavior graph G=(V, E) for training was constructed based on this dataset: nodes V cover types such as commands, users, paths, and system load status, and edges E represent the relationships between nodes such as "preceding execution," "dependency," and "same context." At the same time, initial embedding vectors were generated according to the differences in node types. For example, command nodes were extracted using a pre-trained word2vec or BERT model and mapped to a fixed dimension using MLP. User nodes were converted into 32-dimensional vectors through permission level mapping. Path nodes were generated into corresponding dimension vectors through hierarchical embedding, and system load nodes were generated through numerical standardization. All processed graph data were stored in a graph database and periodically compressed into graph tensor format as training input data for the risk probability prediction model.

[0089] During the network structure initialization and parameter configuration phase, based on the real-time requirements of the operation and maintenance command risk analysis scenario, a lightweight graph neural network architecture was selected, and the number of network layers L was set; the trainable weight matrix W of each layer was initialized. (l) We choose ReLU or GELU as the activation function σ, and define the normalization method for aggregating neighbor node information. Furthermore, we associate the output of the lightweight graph neural network architecture with the subsequent MLP classifier, and initialize the classifier weight matrix W. clf Together with the bias term b, determine the Sigmoid activation function for calculating the risk probability, and the risk level classification threshold.

[0090] During the iterative training and parameter optimization phases, preprocessed graph tensor data is fed into the GNN model in batches. In each training round, node information is first transferred and aggregated through multi-layer convolutional operations of the GNN, and the final embedding vector of the command node is obtained after L layers of iteration. This vector is then fed into the MLP classifier to calculate the risk probability p of the command, which is compared with the actual risk results labeled manually. The prediction bias is calculated using the cross-entropy loss function. Subsequently, based on the gradient descent optimizer, backpropagation is performed along the negative gradient direction of the loss function to synchronously update the weight matrices of each layer of the GNN and the W of the MLP classifier. clf b) Minimize the batch loss. During training, the model performance needs to be evaluated periodically using the validation set, with prediction accuracy measured by the sum of squared residuals. When the validation set loss no longer decreases for several consecutive rounds or reaches the preset number of training rounds, stop the initial training and save the model parameters at this point as the base model.

[0091] During the continuous iterative training phase, after the risk prediction model is deployed, the actual consequences of executed commands are tracked in real time and compared with the model's prediction results (risk probability p). False positives and false negatives are identified, forming incremental training samples. These samples are then used to construct new operational behavior graphs and initial embedding vectors according to preprocessing rules and added to the incremental learning queue. An exponentially weighted moving average algorithm is used to adjust the ratio of new and old training data, giving new samples higher weights while retaining the influence of historical effective data. The mixed dataset is then fed into the saved basic GNN model for small-batch parameter fine-tuning: the GNN layers are retrained to adapt to the evolution of the graph structure in new scenarios, the generation strategy for node initialization vectors is adjusted, and the weights of the MLP classifier are optimized to ensure that the model can dynamically adapt to changes in the operational environment and continuously improve the accuracy of risk prediction.

[0092] The sum of squared residuals is used to measure the accuracy of the risk probability prediction model in predicting the risk probability of operation and maintenance commands. The sum of squared residuals (SSE) is shown in Equation (4):

[0093] (4)

[0094] In formula (4), p i y represents the risk probability predicted by the model. i The actual result is represented by 1 (1 indicates a failure, 0 indicates no failure), a smaller SSE indicates a more accurate prediction, and n represents the number of executions.

[0095] S204. Determine the security control strategy for the operation and maintenance commands to be executed based on the risk probability.

[0096] In some embodiments, after step S203 is executed, the method further includes: determining the risk level of the operation and maintenance command to be executed based on the risk probability score. The risk level can be high risk, medium risk, or low risk. For example, p < 0.3 is determined to be low risk, 0.3 ≤ p < 0.7 is determined to be medium risk, and p ≥ 0.7 is determined to be high risk. Converting the abstract risk probability into a perceptible risk level allows the system to match control and security policies of different strengths according to the severity of the risk.

[0097] Security control strategies include triggering mandatory pop-up prompts, freezing commands, requiring secondary confirmation or approval, demanding dual approval, enabling semantic interpretation, and mandating delayed confirmation for high-risk operations, thereby blocking dangerous operations. For medium-risk operations, rollback suggestions are provided and manual confirmation is required, offering a remedial path for operational errors to balance security and efficiency. For low-risk operations, direct access is allowed and logs are recorded to ensure operational efficiency. This tiered mechanism achieves a dynamic balance between operational security and efficiency. The tiered security control strategy develops differentiated intervention measures for different risk levels, improving the controllability and security of high-risk operations.

[0098] If the risk level is the preset risk level, the system matches the alternative command to be executed from the command library and starts executing the alternative command; if the risk level is not the preset risk level, the system starts executing the operation and maintenance command to be executed.

[0099] Alternative commands are those with similar semantics to the operation and maintenance command to be executed and which have a history of successful operation, or commands with similar semantics to the operation and maintenance command to be executed but with a lower risk level. The command library is automatically accumulated by the system and built by manual maintenance to ensure the availability and security of alternative commands.

[0100] For example, taking the line `rm -rf / var / log` as an example, if its risk probability score is greater than or equal to 0.7, and its risk level is high, the command library will match low-risk alternative commands: `find / var / log-typef-name"*.log"-delete` (delete only the file); `mv / var / log / backup / log_$(date)` (backup first). The chosen alternative command is then `find / var / log-typef-name"*.log"–delete`. This preserves the core business intent of cleaning up logs while mitigating the risk of data loss by deleting only the file and backing up before proceeding.

[0101] The operation and maintenance command processing method provided in this application embodiment further includes: performing semantic parsing on the operation and maintenance command to be executed to obtain the structural semantic unit of the operation and maintenance command to be executed; inputting the structured semantic unit into a language model to obtain the prompt statement output by the language model.

[0102] Among them, structured semantic units include, but are not limited to: operation objects, operation behaviors, and scope of application.

[0103] For example, taking the line `rm -rf / var / log` as an example, if its risk probability score is greater than or equal to 0.7 and its risk level is high, the structured semantic units of the current command are first parsed based on regular expressions, command tree syntax rules, and lexical analysis. These units include: ["Operation type": "delete", "Target": " / var / log", "Scope": "recursive + forced", "Target type": "system log directory"]. This breakdown directly exposes the core risk points of the command. The scope of recursion and forced execution means that the operation is irreversible, and the target type of the system log directory clearly indicates that the operation may lead to log loss and difficulty in tracing the problem.

[0104] When the risk level of the maintenance command to be executed is medium / high, the command is first parsed into structured semantic units. These structured semantic units are then used as input to a language model, which outputs a prompt statement, such as "This command may result in the permanent loss of system logs; a backup is recommended." This prompt clearly indicates the risk consequences and corresponding countermeasures, ensuring that maintenance personnel are clearly aware of the potential harm of the operation and providing a practical security solution.

[0105] In the above embodiments, the extraction of structured semantic units breaks down the ambiguity of command text, enabling precise decomposition of the risk essence of commands and helping operations and maintenance personnel quickly grasp the core risks. Natural language prompts generated based on language models are transformed into intuitive and easy-to-understand safety guidelines, improving the effectiveness of risk intervention. This mechanism clarifies the operational essence by parsing the command structure, informs users of risks and coping methods through prompts, and guides them to choose safer operating methods, effectively lowering the operational threshold and providing reliable operational guidance for novice or temporary operations and maintenance personnel.

[0106] The method for processing operation and maintenance commands provided in this application embodiment further includes: obtaining the execution result of the operation and maintenance command to be executed, and then adjusting the model parameters of the risk prediction model based on the execution result.

[0107] After the execution of the operation and maintenance commands to be analyzed, the execution results are tracked to determine whether the execution caused any problems. If a command predicted to be high-risk (p=0.85) does not actually cause a failure after execution, the system records a false alarm, fine-tunes the model parameters of the risk prediction model, and reduces the sensitivity to this type of command; if a low-risk command (p=0.15) causes a system crash, the system increases the risk probability of similar commands and simultaneously enters the manual review queue.

[0108] In summary, this application provides a method for processing operation and maintenance commands. This method transforms operation and maintenance commands into nodes in an operation behavior graph. By recording the relationships between metadata in the graph, it reduces reliance on human experience. At the same time, structured graph analysis avoids problems such as configuration errors and improper command order. Based on a graph neural network-based risk prediction model, it can utilize the information of nodes (operation and maintenance command metadata) and edges (relationships) in the graph to accurately calculate the risk probability of commands to be executed in dynamic and complex environments, breaking through the limitations of traditional static rules. By determining security control strategies through risk probability, it can provide real-time risk assessment and intervention suggestions before command execution, effectively preventing problems such as system interruption and data anomalies, and adapting to dynamic operation and maintenance environments with multiple platforms, multiple roles, and multiple tasks running in parallel.

[0109] like Figure 3 As shown in the embodiment of this application, another method for processing operation and maintenance commands is provided. This method starts by starting the operation data modeling module, first collecting various operation and maintenance operation-related data, then converting structured operation events into nodes, initializing different types of nodes as vectors, building the relationship between nodes, storing them in the graph database, and then starting the risk assessment module to provide support for subsequent risk judgment.

[0110] Then, the newly input operation and maintenance command is embedded at the node level. After L layers of convolution, the embedding vector is obtained. The node risk value is calculated by a multilayer perceptron classifier. The auxiliary decision-making module is started to receive the risk value predicted by the risk assessment module, and then it is determined whether it is a medium-to-high risk task based on the risk value.

[0111] For medium- to high-risk tasks, alternative commands will be generated based on semantic matching and other methods, and prompts will be generated based on templates and language models. Approval or secondary confirmation operations can be performed based on the prompts, and the execution control module will be activated to determine whether to prohibit execution.

[0112] For tasks that are not of medium or high risk, the execution control module is directly activated to determine whether execution should be prohibited. If the command is executed, the automatic feedback execution and optimization module is entered after the command is executed to determine whether the command execution meets the expected results. If not, the accuracy of the model prediction (false positive rate / false negative rate) is calculated. If the accuracy is lower than the threshold, the feedback model is optimized and adjusted; if the expected results are achieved or the accuracy meets the standard, the process ends.

[0113] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0114] like Figure 4 As shown, embodiments of this application also provide a processing device for operation and maintenance commands, the device comprising:

[0115] Module 401 is used to obtain the operation and maintenance commands to be executed.

[0116] The graph index module 402 is used to determine the target metadata corresponding to the operation and maintenance command to be executed based on the pre-built operation behavior graph; wherein, the operation behavior graph uses the metadata of the operation behavior events of the operation and maintenance command as nodes and the association relationship between the metadata as edges;

[0117] The risk probability prediction module 403 is used to obtain the risk probability of the operation and maintenance command to be executed output by the risk prediction model based on the target metadata and the risk prediction model; wherein, the risk prediction model is built based on a graph neural network and is trained by sample commands and their risk probabilities;

[0118] The control strategy determination module 404 is used to determine the security control strategy for the operation and maintenance commands to be executed based on the risk probability.

[0119] As an optional implementation method provided in this application, the risk prediction model includes an embedding layer and a multilayer perceptron;

[0120] The risk probability prediction module 403 is specifically used to: perform node-level embedding calculations on the target metadata through the embedding layer to obtain the embedding vector of the target metadata; and perform probability calculations on the embedding vector through the multilayer perceptron to obtain the risk probability of the operation and maintenance command to be executed.

[0121] As an optional implementation provided in this application embodiment, the risk probability prediction module 403, when performing node-level embedding calculation on the target metadata through the embedding layer to obtain the embedding vector of the target metadata, specifically performs the following: obtaining the neighbor node vectors of the target metadata in the adjacent upper layer through the embedding layer; performing weight transformation on each neighbor node vector according to the weight matrix of the adjacent upper layer; normalizing and adding the weighted neighbor node vectors to obtain the sum result; and obtaining the embedding vector of the target metadata based on the sum result and the activation function.

[0122] As an optional implementation provided in this application, the risk probability prediction module 403 is further configured to: determine the risk level of the operation and maintenance command to be executed based on the risk probability; if the risk level is a preset risk level, match an alternative command for the operation and maintenance command to be executed from the command library to start executing the alternative command; if the risk level is not a preset risk level, start executing the operation and maintenance command to be executed.

[0123] As an optional implementation provided in this application, the device further includes a prompting module, used for: performing semantic parsing on the operation and maintenance command to be executed to obtain a structured semantic unit of the operation and maintenance command to be executed; inputting the structured semantic unit into a language model to obtain a prompting statement output by the language model.

[0124] As an optional implementation provided in this application, the device further includes a model optimization module, used for: obtaining the execution result of the operation and maintenance command to be executed; and adjusting the model parameters of the risk prediction model according to the execution result.

[0125] As an optional implementation method provided in this application, the metadata of the operation behavior of the operation and maintenance command includes at least one of the following: system load status, command content, execution path, user identity, preceding command sequence, and host identifier.

[0126] As an optional implementation provided in this application, the device further includes a graph construction module, used to obtain metadata and context information of the operation behavior events of operation and maintenance commands; construct structured operation behavior events based on the metadata; and construct an operation behavior graph based on the structured operation behavior events and context information.

[0127] As an optional implementation provided in this application, the map index module 402 is further configured to: if the target metadata is command content, extract a text vector from the command content and use the text vector as input to the risk prediction model; if the target metadata is user identity, determine the user identity's permission level and use the permission level as input to the risk prediction model.

[0128] As an optional implementation provided in this application, the graph index module 402 is further configured to: if the target metadata is an execution path, split the execution path into multiple sub-units; perform an embedding operation on each sub-unit to obtain the embedding vector of each sub-unit; average the embedding vectors of each sub-unit to obtain the path layer feature vector of the path node, so as to use the path layer feature vector as the input of the risk prediction model.

[0129] As an optional implementation provided in this application, the map index module 402 is further configured to: if the target metadata is the system load status, obtain the resource network core indicators corresponding to the system load status; the resource network core indicators include at least one of processor load, memory occupancy rate and network status; and concatenate the standardized resource network core indicators to obtain the system load status vector, so as to use the system load status vector as the input of the risk prediction model.

[0130] For a description of the features in the embodiment corresponding to the operation and maintenance command processing device, please refer to the relevant description in the embodiment corresponding to the operation and maintenance command processing method, which will not be repeated here.

[0131] like Figure 5 As shown, embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above-described operation and maintenance command processing method embodiments.

[0132] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described operation and maintenance command processing method embodiments when running.

[0133] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0134] The embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in the processing method embodiments of any of the above-described operation and maintenance commands.

[0135] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in the above-described embodiments of the operation and maintenance command processing method.

[0136] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0137] The foregoing has provided a detailed description of the method, apparatus, device, storage medium, and program product for processing operation and maintenance commands provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only intended to aid in understanding the method and core ideas of this application. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A method for processing operation and maintenance commands, characterized in that, include: Retrieve operation and maintenance commands to be executed; Based on a pre-constructed operation behavior graph, the target metadata corresponding to the operation and maintenance command to be executed is determined; wherein, the operation behavior graph uses the metadata of the operation behavior events of the operation and maintenance command as nodes, and the association relationship between the metadata as edges; Based on the target metadata and the risk prediction model, the risk probability of the operation and maintenance command to be executed is obtained from the output of the risk prediction model; wherein, the risk prediction model is constructed based on a graph neural network and is trained from sample commands and their risk probabilities; The security control strategy for the operation and maintenance command to be executed is determined based on the risk probability. After determining the target metadata corresponding to the operation and maintenance command to be executed based on the pre-constructed operation behavior graph, and before obtaining the risk probability of the operation and maintenance command to be executed output by the risk prediction model based on the target metadata and the risk prediction model, the method further includes: If the target metadata is command content, then a pre-trained word2vec model or a Transformer-based bidirectional encoder BERT model is used to extract text vectors from the command content, and the text vectors are used as input to the risk prediction model. If the target metadata is a user identity, then the permission level of the user identity is determined by one-hot encoding or permission level mapping, so as to use the permission level as the input of the risk prediction model; If the target metadata is an execution path, then the execution path is split into multiple sub-units; an embedding operation is performed on each sub-unit to obtain the embedding vector of each sub-unit; the embedding vectors of each sub-unit are averaged to obtain the path hierarchical feature vector of the execution path, and the path hierarchical feature vector is used as the input of the risk prediction model. If the target metadata is the system load status, then the resource network core indicators corresponding to the system load status are obtained; the resource network core indicators include at least one of processor load, memory utilization, and network status; the standardized resource network core indicators are concatenated to obtain the system load status vector, which is used as the input of the risk prediction model; The risk prediction model includes an embedding layer and a multilayer perceptron; The step of obtaining the risk probability of the operation and maintenance command to be executed, output by the risk prediction model, based on the target metadata and the risk prediction model includes: obtaining the neighbor node vectors of the target metadata in the adjacent upper layer through the embedding layer; performing weight transformation on each neighbor node vector according to the weight matrix of the adjacent upper layer, wherein the weight matrix automatically learns the risk contribution weights of different neighbor nodes to the current node during training; performing normalization processing based on node degree on the weighted neighbor node vectors and adding them to obtain the sum result; obtaining the embedding vector of the target metadata based on the sum result and the activation function; and performing nonlinear transformation and mapping on the embedding vector through the multilayer perceptron to obtain the risk probability of the operation and maintenance command to be executed.

2. The method according to claim 1, characterized in that, After obtaining the risk probability of the operation and maintenance command to be executed as output by the risk prediction model based on the target metadata and the risk prediction model, the method further includes: The risk level of the operation and maintenance command to be executed is determined based on the risk probability. If the risk level is a preset risk level, a replacement command for the operation and maintenance command to be executed is matched from the command library to initiate the execution of the replacement command; If the risk level is not the preset risk level, the pending operation and maintenance command will be executed.

3. The method according to claim 1, characterized in that, The method further includes: The operation and maintenance command to be executed is semantically parsed to obtain the structured semantic unit of the operation and maintenance command to be executed; The structured semantic unit is input into the language model to obtain the prompt statement output by the language model.

4. The method according to claim 1, characterized in that, The method further includes: Obtain the execution result of the operation and maintenance command to be executed; Adjust the model parameters of the risk prediction model based on the execution results.

5. The method according to claim 1, characterized in that, The process of constructing the operational behavior map includes: Obtain the metadata and context information of the operation and maintenance command's behavior events; Construct structured operational behavior events based on the aforementioned metadata; The operation behavior graph is constructed based on the structured operation behavior events and the context information.

6. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the processing method for the operation and maintenance commands as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Intelligent operation and maintenance management method and system based on knowledge graph

    CN116611813A

  • Operation and maintenance operation risk management and control method and device and electronic equipment

    CN117196485A