Method and device for detecting cloud network equipment and nonvolatile storage medium

By processing the multimodal heterogeneous data of cloud network devices through anomaly probability analysis models and dynamic pruning technology, the problem of low accuracy in anomaly detection of cloud network devices is solved, efficient anomaly detection and real-time alarms are achieved, and the security and stability of the cloud network environment are ensured.

CN120768792APending Publication Date: 2025-10-10CHINA TELECOM CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510921052.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

In the cloud computing and network converged environment, it is difficult for anomaly detection of cloud network devices to effectively integrate multimodal heterogeneous data, resulting in low accuracy of detection results.

Method used

Anomaly detection is achieved by using an anomaly probability analysis model to process and analyze the multimodal feature fusion vector. By generating a multimodal feature fusion vector and a dynamic association graph, and combining knowledge distillation and dynamic pruning techniques to train a neural network model, the dynamic dependency relationship between cloud network devices is captured.

Benefits of technology

It improves the accuracy of anomaly detection of cloud network equipment, enhances the real-time and security stability of detection results, and ensures the safe operation of the cloud network environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120768792A_ABST
    Figure CN120768792A_ABST
Patent Text Reader

Abstract

The invention discloses a method and device for detecting cloud network equipment and a nonvolatile storage medium. The method comprises the following steps: acquiring various operation data of to-be-detected equipment in a detection time period, and generating a multi-modal feature fusion vector according to the various operation data; for each to-be-detected device, an anomaly probability analysis model is adopted to process and analyze the multi-modal feature fusion vector, a probability matrix output by the anomaly probability analysis model is obtained, and each element in the probability matrix is used for indicating the probability that the to-be-detected device has one type of anomaly; according to a plurality of probability matrixes corresponding to a preset time window, whether the to-be-detected device is abnormal in the preset time window is judged, a judgment result is obtained, the preset time window comprises a plurality of continuous detection time periods, and the judgment result records the abnormal type of the to-be-detected device; and determining the alarm level of each abnormal type according to a plurality of judgment results corresponding to the plurality of preset time windows.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of network security technology, and more specifically, to a method and apparatus for detecting cloud network equipment, and a non-volatile storage medium. Background Art

[0002] In the complex environment of cloud computing and network integration, anomaly detection of cloud-network devices (such as servers, switches, and load balancers) is a core task to ensure system stability. Cloud-network devices typically involve complex network structures, large amounts of log data, and various types of attack modes. Related technologies have been struggling to integrate heterogeneous data generated by cloud-network devices, such as structured logs, performance metrics, network traffic, and unstructured text logs, when performing anomaly detection on cloud-network devices. This creates difficulties in processing multimodal, heterogeneous data, resulting in low accuracy in detection results.

[0003] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention

[0004] The embodiments of the present application provide a method and apparatus for detecting cloud network devices, and a non-volatile storage medium, to at least solve the technical problem of low accuracy of detection results when detecting whether a cloud network device is abnormal, caused by the difficulty of the related art method for detecting whether a cloud network device is abnormal in processing multimodal heterogeneous data.

[0005] According to one aspect of an embodiment of the present application, a method for detecting cloud network devices is provided, including: obtaining multiple operating data of the device to be detected during a detection period, and generating a multimodal feature fusion vector based on the multiple operating data, wherein the device to be detected is a cloud network device in a cloud network environment, and the number of the devices to be detected is multiple; for each device to be detected, using an abnormality probability analysis model to process and analyze the multimodal feature fusion vector to obtain a probability matrix output by the abnormality probability analysis model, wherein each element in the probability matrix is ​​used to indicate the probability of a class of abnormality existing in the device to be detected, and the abnormality probability analysis model is an initial neural network model with multiple neural network layers deleted during the training process; judging whether the device to be detected has an abnormality in the preset time window according to multiple probability matrices corresponding to the preset time window, and obtaining a judgment result, wherein the preset time window contains multiple continuous detection periods, and the judgment result records the abnormality type of the device to be detected; determining the alarm level of each abnormality type according to the multiple judgment results corresponding to the multiple preset time windows.

[0006] Optionally, a multimodal feature fusion vector is generated based on multiple types of operating data, including: generating a dynamic association graph based on multiple types of operating data, wherein the types of operating data include: log data, indicator data and traffic data, each node in the dynamic association graph represents a device to be detected, and the edges of the dynamic association graph are constructed based on the spatiotemporal association relationship and semantic association relationship between the devices to be detected; performing feature extraction on the dynamic association graph to obtain a high-order feature vector of the cloud network device, wherein the high-order feature vector is used to describe multi-dimensional information of multiple devices to be detected, including: operating information of the devices to be detected, and the association relationship between multiple devices to be detected; generating a multimodal feature fusion vector based on the high-order feature vector and the feature vector obtained by feature extraction of each type of operating data.

[0007] Optionally, feature extraction is performed on the dynamic association graph to obtain a high-order feature vector of the cloud network device, including: for each edge in the dynamic association graph, determining the two nodes associated with the edge, and obtaining two sets of operating data corresponding to the two nodes; determining the semantic similarity of events occurring in the two devices to be detected corresponding to the two nodes within the detection period based on the two sets of operating data; determining an event with a semantic similarity greater than a preset semantic similarity as a target event, and determining the target time when the target event occurs in the two devices to be detected corresponding to the two nodes within the detection period based on the two sets of operating data; determining the spatiotemporal correlation strength of the two devices to be detected corresponding to the two nodes based on the target time; determining the weight of the edge based on the spatiotemporal correlation strength and the semantic similarity; determining the adjacency matrix based on multiple weights corresponding to multiple edges; and performing convolution processing on the adjacency matrix to obtain a high-order feature vector.

[0008] Optionally, an abnormal probability analysis model is used to process and analyze the multimodal feature fusion vector to obtain a probability matrix output by the abnormal probability analysis model, including: determining multiple position coding vectors, and fusing the multiple position coding vectors and multiple multimodal feature fusion vectors to obtain a feature sequence, wherein the feature sequence is composed of multiple time series feature vectors, each time series feature vector is generated based on a position coding vector and a multimodal feature vector, each position coding vector is used to indicate the relative position of the corresponding time series feature vector in the feature sequence, and the position coding vector is generated during the training process of the abnormal probability analysis model; determining the processing order according to the multiple relative positions, and processing each time series feature vector in the feature sequence in turn according to the processing order in the probability prediction layer until all the time series feature vectors are processed to obtain a hidden layer feature vector, wherein when processing the next time series feature vector, the processing result of the previous time series feature vector and the next time series feature vector are used together as the probability prediction layer input data; performing dimension mapping and normalization on the hidden layer feature vector to obtain a probability matrix.

[0009] Optionally, the abnormal probability analysis model is trained by the following method: obtaining a training data set, wherein the training data set includes: historical operation data and multiple real labels, the historical operation data is multiple types of operation data generated by multiple IoT devices in the historical operation period, and the real label is used to indicate the operation status of the IoT device in the historical operation period, and the operation status includes: normal operation and abnormal operation; performing feature fusion processing on the historical operation data to obtain a historical multimodal feature fusion vector; inputting the historical multimodal feature fusion vector into the teacher model and the initial neural network model respectively to obtain the prediction result output by the initial neural network model and the soft label output by the teacher model, wherein the teacher model is a neural network model with a larger number of parameters than the initial neural network model, the soft label is the operation status of the IoT device in the historical operation period predicted by the teacher model, and the prediction result is the operation status of the IoT device in the historical operation period predicted by the initial neural network model; determining a first loss function based on the soft label and the prediction result, determining a second loss function based on the prediction result and the real label, and determining a total loss function based on the first loss function and the second loss function; determining the training process of the initial neural network model based on the total loss function, wherein the training process includes: continuing iterative training and stopping iterative training.

[0010] Optionally, during the training process, multiple neural network layers of the initial neural network model are deleted by the following method: obtaining the absolute value of the weight of each layer in the initial neural network model, and arranging the multiple absolute values ​​into a weight sequence in ascending order; determining a pruning threshold based on the weight sequence, wherein the pruning threshold is used to assist in determining the deleted neural network layer; assigning an absolute value less than the pruning threshold to an invalid value, and assigning an absolute value greater than the pruning threshold to a valid value; determining the neural network layer corresponding to the invalid value as the deleted neural network layer, until the sparsity of the initial neural network model reaches a preset sparsity, stopping the iteration, and obtaining an abnormal probability analysis model, wherein the preset sparsity is used to indicate the ratio of the number of deleted neural network layers to the number of all neural network layers contained in the initial neural network model.

[0011] Optionally, determining a pruning threshold based on the weight sequence includes: during the first iterative training, determining the deletion ratio corresponding to each absolute value, and when the deletion ratio is equal to the preset sparsity, determining the absolute value corresponding to the deletion ratio as the pruning threshold, wherein the deletion ratio is used to indicate the ratio of the number of other absolute values ​​in the weight sequence that are smaller than the absolute value to the total number of absolute values ​​contained in the weight sequence; during non-first iterative training, updating the weight sequence according to the results of the previous iterative training to obtain a new weight sequence; determining the deletion ratio corresponding to each absolute value in the new weight sequence, and when the deletion ratio is equal to the preset sparsity, determining the absolute value corresponding to the deletion ratio as the current pruning threshold; and determining the pruning threshold based on the historical pruning threshold and the current pruning threshold used in the previous iterative training process.

[0012] Optionally, whether the device to be detected has an abnormality is judged based on multiple probability matrices corresponding to a preset time window to obtain a judgment result, including: for each probability matrix, determining a target element in the probability matrix that is greater than a preset probability value, and when the target element exists in any probability matrix, determining the judgment result as that the device to be detected has an abnormality, and determining the abnormality type corresponding to the target element as the abnormality type of the device to be detected, and the abnormality types include: network attack, performance bottleneck, hardware failure; when the target element does not exist in multiple probability matrices corresponding to the preset time window, determining the judgment result as that the device to be detected has no abnormality.

[0013] Optionally, the alarm level of each abnormality type is determined based on multiple judgment results corresponding to multiple preset time windows: the number of alarms corresponding to each abnormality type is determined, wherein the number of alarms is the number of times the abnormality type is recorded in multiple judgment results; the alarm level of the abnormality type corresponding to the number of alarms greater than or equal to the first preset value is determined as the first level; the alarm level of the abnormality type corresponding to the number of alarms greater than or equal to the second preset value is determined as the second level, wherein the second preset value is greater than the first preset value, and the second level is higher than the first level; the alarm level of the abnormality type corresponding to the number of alarms greater than or equal to the third preset value is determined as the third level, wherein the third preset value is greater than the second preset value, and the third level is higher than the second level.

[0014] According to another aspect of an embodiment of the present application, a device for detecting cloud network devices is also provided, including: an acquisition module for acquiring multiple operating data of the device to be detected during the detection period, and generating a multimodal feature fusion vector based on the multiple operating data, wherein the device to be detected is a cloud network device in a cloud network environment, and the number of devices to be detected is multiple; a processing module for processing and analyzing the multimodal feature fusion vector for each device to be detected using an abnormality probability analysis model to obtain a probability matrix output by the abnormality probability analysis model, wherein each element in the probability matrix is ​​used to indicate the probability of a class of abnormality existing in the device to be detected, and the abnormality probability analysis model is an initial neural network model in which multiple neural network layers are deleted during the training process; a judgment module for judging whether the device to be detected has an abnormality in the preset time window according to multiple probability matrices corresponding to the preset time window, and obtaining a judgment result, wherein the preset time window contains multiple continuous detection periods, and the judgment result records the abnormality type of the device to be detected; a determination module for determining the alarm level of each abnormality type according to multiple judgment results corresponding to multiple preset time windows.

[0015] According to another aspect of an embodiment of the present application, a non-volatile storage medium is further provided, in which a computer program is stored. The device where the non-volatile storage medium is located executes the above-mentioned method for detecting cloud network devices by running the computer program.

[0016] According to another aspect of an embodiment of the present application, an electronic device is also provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute the above-mentioned method for detecting cloud network devices through the computer program.

[0017] According to another aspect of an embodiment of the present application, a computer program product is also provided, comprising computer instructions, which, when executed by a processor, implement the steps of the above-mentioned method for detecting cloud network devices.

[0018] In an embodiment of the present application, a plurality of operating data of a device to be detected during a detection period is obtained, and a multimodal feature fusion vector is generated based on the plurality of operating data, wherein the device to be detected is a cloud network device in a cloud network environment, and the number of devices to be detected is multiple; for each device to be detected, an abnormal probability analysis model is used to process and analyze the multimodal feature fusion vector to obtain a probability matrix output by the abnormal probability analysis model, wherein each element in the probability matrix is ​​used to indicate the probability that a class of abnormalities exists in the device to be detected, and the abnormal probability analysis model is an initial neural network model in which multiple neural network layers are deleted during the training process; based on multiple probability matrices corresponding to a preset time window, it is judged whether the device to be detected has an abnormality in a preset time window to obtain a judgment result, wherein the preset time window contains multiple continuous detection periods, and the judgment result records the abnormality type of the device to be detected; a method for determining the alarm level of each abnormality type based on multiple judgment results corresponding to multiple preset time windows The method introduces an anomaly probability analysis model trained using knowledge distillation and dynamic pruning technology, and uses the anomaly probability analysis model to process multimodal heterogeneous data of cloud network devices to realize anomaly detection of cloud network devices, thereby solving the problem that the existing technology cannot process multimodal heterogeneous data. When there are multiple cloud network devices to be detected, the anomaly probability analysis model can capture the dynamic dependency relationship between cloud network devices, and detect abnormal cloud network devices in the same cloud network environment based on the dynamic dependency relationship, thereby achieving the purpose of improving detection accuracy, thereby achieving the technical effect of improving the accuracy of detection results of detecting whether cloud network devices are abnormal. Furthermore, the alarm level is determined according to the prediction result output by the anomaly probability analysis model, which improves the real-time nature of the alarm and effectively ensures the safe and stable operation of the cloud network environment, thereby solving the technical problem of low accuracy of detection results when detecting whether cloud network devices are abnormal due to the fact that the method for detecting whether cloud network devices are abnormal in related technologies is difficult to process multimodal heterogeneous data. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0020] Figure 1 This is a hardware structure block diagram of a computer terminal for implementing a method for detecting cloud network devices according to an embodiment of the present application;

[0021] Figure 2 This is a flowchart of the steps of a method for detecting cloud network devices according to an embodiment of the present application;

[0022] Figure 3 is a schematic diagram of a data processing process of an abnormal probability analysis model according to an embodiment of the present application;

[0023] Figure 4 is a model structure schematic diagram of a teacher model according to an embodiment of the application;

[0024] Figure 5 is a pseudo code of a dynamic pruning algorithm according to an embodiment of the application;

[0025] Figure 6 is a structural diagram of an apparatus for detecting a cloud network device according to an embodiment of the application;

[0026] Figure 7 is a work flow diagram of an apparatus for detecting a cloud network device according to an embodiment of the application;

[0027] Figure 8 is an execution framework schematic diagram of a method for detecting a cloud network device according to an embodiment of the application. DETAILED DESCRIPTION

[0028] In order to enable persons skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by persons skilled in the art without creative work should fall within the scope of protection of the present application.

[0029] It should be noted that the terms "first", "second", and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0030] In the related art, when detecting the cloud network device, it is difficult to fuse the heterogeneous data such as structured logs, performance indicators, network traffic, and unstructured text logs generated by the cloud network device, and there is a problem of difficulty in processing multi-modal heterogeneous data. In addition, the method for detecting the cloud network device in the related art also has the following problems: 1) The model used is mostly based on a static graph or a fixed time window, and cannot capture the dynamic association relationship between the cloud network devices. In the actual running process, the dependency relationship between the cloud network devices changes in real time with the load, or the network topology changes due to fault switching or traffic scheduling. Therefore, the related art cannot capture the dynamic association relationship between the cloud network devices, resulting in the inability to timely perceive the sudden abnormality. 2) The model structure is complex, and large-scale data processing is required during prediction, which takes a long time to predict and cannot meet the requirements of low-latency abnormality detection in the cloud environment. In order to solve the above problems, the related solutions are provided in the embodiments of the present application, which are described in detail below.

[0031] According to the embodiments of the present application, a method for detecting a cloud network device is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0032] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 A hardware structure block diagram of a computer terminal for implementing the method of detecting a cloud network device is shown. As Figure 1 shown, the computer terminal 10 can include one or more processors 102 (the processor 102 can include but is not limited to a microprocessor MCU or a programmable logic device FPGA processing device), a memory 104 for storing data, and a transmission device 106 for communication function. In addition, it can also include a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which can be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. Those skilled in the art can understand that Figure 1 The structure shown is only schematic, which does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 can include more or fewer components than those shown in Figure 1 or have a different configuration from Figure 1 shown.

[0033] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry." The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuitry may be a single, independent processing module, or may be incorporated in whole or in part into any of the other components of the computer terminal 10. As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).

[0034] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the method for detecting cloud network devices in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, realizing the above-mentioned method for detecting cloud network devices. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0035] The transmission device 106 is configured to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of the computer terminal 10. In one embodiment, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device 106 may be a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.

[0036] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 .

[0037] The embodiment of the present application provides a method for detecting cloud network devices that can be run in the above operating environment. Figure 2 This is a flowchart of the steps of the method for detecting cloud network devices provided in an embodiment of the present application. Figure 2 As shown, the method includes the following steps:

[0038] Step S202: Acquire multiple operating data of the device to be detected during the detection period, and generate a multimodal feature fusion vector based on the multiple operating data, wherein the device to be detected is a cloud network device in a cloud network environment, and the number of the devices to be detected is multiple.

[0039] The detection method for cloud network devices provided in the embodiment of the present application obtains the operating data of the device to be detected during the detection period in real time, and determines whether there is any abnormality in the operation of the device to be detected during the detection period by fusing the obtained operating data; the above-mentioned operating data obtained include: log data, indicator data, network traffic data and other types of operating data, among which the log data records the operation records and events of the cloud network device (such as failure to connect to the server using the Secure Shell Protocol (SSH), abnormal port closure), etc.; the indicator data records the real-time performance indicators of the cloud network device, including: central processing unit (CPU) utilization, memory usage, bandwidth utilization, etc.; the traffic data includes the source Internet Protocol address (IP), port, protocol, packet size, latency, etc. of the cloud network device. The detection of cloud network devices provided in the embodiments of the present application can be applied to one cloud network device in a cloud network environment or to multiple cloud network devices in a cloud network environment at the same time. Therefore, the device to be detected can be one cloud network device or multiple cloud network devices. Cloud network devices refer to devices running in a cloud network environment, such as servers, switches, load balancers, etc.; cloud network environment refers to the operating environment provided by the infrastructure that integrates cloud computing and network technology. In a cloud network environment, resources can be flexibly allocated and scheduled between the cloud and the network, and data and applications can be seamlessly migrated between cloud resources and the network. In step S202, when the operating data generated by the device to be detected during the detection period is fused, a multimodal feature fusion vector that integrates the features of each type of operating data is first generated based on the multiple types of operating data. When the multiple types of operating data are fused to generate the multimodal feature fusion vector, the multiple types of operating data generated by the same device to be detected at the same time are fused.

[0040] According to some optional embodiments of the present application, a multimodal feature fusion vector is generated based on multiple types of operating data, including: generating a dynamic association graph based on multiple types of operating data, wherein the types of operating data include: log data, indicator data and traffic data, each node in the dynamic association graph represents a device to be detected, and the edges of the dynamic association graph are constructed based on the spatiotemporal association relationship and semantic association relationship between the devices to be detected; performing feature extraction on the dynamic association graph to obtain a high-order feature vector of the cloud network device, wherein the high-order feature vector is used to describe multi-dimensional information of multiple devices to be detected, including: operating information of the devices to be detected, and the association relationship between multiple devices to be detected; generating a multimodal feature fusion vector based on the high-order feature vector and the feature vector obtained by feature extraction of each type of operating data.

[0041] In this embodiment, the following method can be used to process various types of operating data into multimodal feature fusion vectors, including the following steps: Step 1. Construct a spatiotemporal-semantic association graph (i.e., a dynamic association graph) G based on various types of operating data. The dynamic association graph G is used to reflect the spatiotemporal and semantic association relationships of multiple cloud network devices. Each node (V) in the dynamic graph represents a device to be detected (i.e., a cloud network device). When there is a semantic association relationship or a spatiotemporal association relationship between two devices to be detected, an edge (E(t)) is generated between the two nodes corresponding to the two devices to be detected. Step 2. Feature extraction is performed on the dynamic association graph to obtain a high-order feature vector, wherein the high-order feature vector is used to record the multi-dimensional information of the device to be detected. The multi-dimensional information includes: the operating information of the device to be detected itself, and the association relationship (such as dependency relationship) between multiple devices to be detected. Step 3: Fuse the feature vector of each type of operation data with the high-order feature vector obtained above to obtain a multimodal feature fusion vector, wherein the feature vector of each type of operation data is obtained by extracting features from each type of operation data. For example, the multiple types of operation data obtained in step S202 are: log data, indicator data, and flow data. Then, the feature vector of each type of operation data includes: log event features obtained by extracting features from log data, indicator time series features obtained by extracting features from indicator data, and flow session features obtained by extracting features from flow data. Since there are multiple data, when performing fusion, the matrix X composed of multiple log event features is combined into log , X is composed of multiple indicator time series features metric and a matrix X consisting of multiple traffic session features flow and a matrix H consisting of multiple high-order eigenvectors GCN to integrate.

[0042] In this embodiment, when executing step 1 to construct a dynamic association graph, the various types of operating data obtained can be first converted into structured data. For example, when the operating data includes log data, indicator data, and traffic data, the following method can be used to convert them into structured data. For log data, regular expressions are used to parse device operation records and events, extract fields such as timestamps, device identifiers (IDs), event levels (information level INFO, warning level WARN, error level ERROR), and generate a log event table. The log event table is the structured data of the log data. For indicator data, indicators such as the device's CPU utilization, memory usage, and network bandwidth usage are normalized and expressed as percentages. For example, the CPU utilization of cloud network device A within a certain timestamp is normalized to 80%, and the memory usage is normalized to 70%. The timestamps are aligned by device ID to generate an indicator time series table. The indicator time series table is the structured data of the indicator data. For traffic data, the five-tuple of traffic data (protocol, source IP, source port, destination IP, destination port), traffic size, and latency are counted to form session-level metadata. Timestamps are aligned by device ID to generate a traffic session table, which is the structured data of the indicator data. Next, when constructing the dynamic association graph G = (V, E(t)), edges are generated by capturing the dynamic dependencies between the devices to be detected. Dynamic dependencies include spatiotemporal and semantic dependencies. Spatiotemporal dependencies refer to the interdependence or interaction between cloud network devices in the temporal and spatial dimensions. For example, the interdependence in the temporal dimension is as follows: if the indicator similarity of two indicator data sequences corresponding to two cloud network devices is greater than the preset indicator similarity, the two cloud network devices are considered to have a spatiotemporal association relationship. For another example, if two cloud network devices are physically or logically adjacent, such as if they are located in the same data center or virtual network and share the same network link, the two cloud network devices are considered to have a spatiotemporal association relationship. For example, if both devices A and B under test experience high CPU loads within one minute, a spatiotemporal relationship is determined between the two devices under test. An edge E(t) is then established between them, with the weight of this edge determined by load similarity. Semantic relationships can be determined through events recorded in log data. For example, a semantic relationship exists between two cloud network devices involved in the same event. When constructing edges based on semantic relationships, edges can be defined based on log event causality or traffic path analysis. For example, if device C's disk failure log is often accompanied by a storage service response timeout on device D, a semantic edge E(t) is established between the two.

[0043] In this embodiment, step 2 of performing feature extraction on the dynamic correlation graph to obtain a high-order feature vector can be implemented using a neural network model. For example, a multi-branch convolutional attention neural network model (DMF-GNN) can be used to process multiple types of running data to generate a multi-modal feature vector. Step 3 of performing feature fusion to generate a multi-modal feature vector can also be implemented using DMF-GNN. The process of DMF-GNN performing feature fusion to generate a multi-modal feature vector can be described using the formula F fusion =∑ k a k ·Z k ,(k=GCN、log、metric、flow) where Z k represents the feature vector of the kth modality, a k is the weight of the kth modality, and the modalities include: log, metric, flow, and high-order feature modality (GCN). a k can be determined according to the following formula: a k =softmax(v T ·tanh(W α ·Z k )),(k=GCN、log、metric、flow) where softmax is an activation function that converts a set of numerical values into a probability distribution, v is a vector used to calculate attention weights (i.e., attention vector), W T is a model parameter of DMF-GNN, v α is the transpose matrix of the attention vector matrix; W α is an attention weight matrix, which is a model parameter of DMF-GNN; and tanh is a hyperbolic tangent function, which is an activation function used to map input values (i.e., W k ·Z GCN ) to the value interval [-1, 1].

[0044] When performing feature fusion in step 3, DMF-GNN first projects each modality feature to the same dimension using a fully connected layer, for example, to 512 dimensions, to achieve feature alignment. The process of feature alignment can be described using the following formula: In the formula, Z log is the linear transformation result (i.e., projection result) of the high-order feature vector, Z metric is the linear transformation result (i.e., projection result) of the log event feature, Z flow is the linear transformation result (i.e., projection result) of the metric time series feature, and Z GCNis the weight matrix of high-order features, W log is the weight matrix of log data, W metric is the weight matrix of indicator data, W flow is the weight matrix of traffic data, b GCN is the bias term of the high-order feature, b log is the bias term of the log data, b metric is the bias term of the indicator data, b flow is the bias term of the traffic data; the above multiple weight matrices and bias terms are the model parameters of DMF-GNN, which are continuously updated during the model training process and remain unchanged after the training is completed.

[0045] Optionally, feature extraction is performed on the dynamic association graph to obtain a high-order feature vector of the cloud network device, including: for each edge in the dynamic association graph, determining the two nodes associated with the edge, and obtaining two sets of operating data corresponding to the two nodes; determining the semantic similarity of events occurring in the two devices to be detected corresponding to the two nodes within the detection period based on the two sets of operating data; determining an event with a semantic similarity greater than a preset semantic similarity as a target event, and determining the target time when the target event occurs in the two devices to be detected corresponding to the two nodes within the detection period based on the two sets of operating data; determining the spatiotemporal correlation strength of the two devices to be detected corresponding to the two nodes based on the target time; determining the weight of the edge based on the spatiotemporal correlation strength and the semantic similarity; determining the adjacency matrix based on multiple weights corresponding to multiple edges; and performing convolution processing on the adjacency matrix to obtain a high-order feature vector.

[0046] According to the above embodiment, DMF-GNN can be used to extract features from the dynamic association graph to obtain a high-order feature vector H GCN In this embodiment, when DMF-GNN extracts high-order features from the dynamic association graph, it takes into account that the status and interaction relationships of cloud network devices change rapidly over time (such as sudden traffic surges, instantaneous surges in CPU load, and other sudden device behaviors). A dynamic adjacency matrix is ​​used to update the weights between nodes in the dynamic graph (i.e., the weights of edges) in real time, capturing time-sensitive and semantically related associations and accurately modeling transient anomalies. The adjacency matrix is ​​a matrix with the weights of the edges in the dynamic association graph as elements. The calculation method of the elements of the dynamic adjacency matrix is ​​as follows: Where A ij represents the strength of the association between node i and node j (i.e., the weight of the edge associating node i and node j), A ij The larger the value of is, the stronger the spatiotemporal correlation strength and semantic similarity between nodes are; ij The value of 0 indicates that node i and node j have no association in the current time window. i Indicates the timestamp of the co-occurrence event of node i in the time window, t jIndicates the timestamp of the co-occurrence event of node j within the time window. t is the time decay rate control parameter, which controls the decay speed of time correlation, σ t The larger the value of , the slower the time decay, allowing for longer time span correlations; t The smaller the value of , the more severe the time decay, and only the events in a very short time are concerned. i ,s j ) represents the semantic similarity between the log or event descriptions of node i and node j, sim(s i ,s j ) can be calculated using the formula sim(s i ,s j )=(s i ·s j ) / (‖s i ‖‖s j ‖) is calculated, where s i is the text vector of the running data of node i, s j Is the text vector of the running data of node j. The dynamic adjacency matrix strengthens the spatiotemporal co-occurrence and semantic association between cloud network devices through the dynamic fusion of time decay function and semantic similarity, weakens historical noise, improves detection accuracy, and only calculates non-zero weights for nodes that co-occur within the time window, sparse dynamic update, reduces the consumption of computing resources, and adapts to the resource limitations of edge devices. In this embodiment, in order to ensure that non-zero weights are calculated only for nodes that co-occur within the time window, the nodes that co-occur within the time window are determined based on the similarity of events that occur during the detection period of the cloud network device corresponding to the node (i.e., the device to be detected). Specifically, when the similarity of events that occur during the detection period of two cloud network devices connected by the same edge is greater than the preset similarity, it is determined that the two nodes have a co-occurrence event (i.e., the target event) during the detection period, and the two nodes are co-occurring nodes within the time window; the events that occur during the detection period of each cloud network device can be extracted from the log data of the cloud network device, and the similarity of the events can be determined using the semantic similarity of the text describing the event; the period during which the co-occurrence event of the two nodes is the time window for applying the above formula. The constructed dynamic association graph can be represented by G(t)=(V,E(t)), where node V represents the cloud network device, E(t) represents the edge connecting the nodes, and its weight A ij (t) Updated by sliding window over time.

[0047] In this embodiment, after determining the matrix with the edges in the dynamic association graph as elements (i.e., the adjacency matrix), the adjacency matrix is ​​convolved to obtain a high-order feature matrix. During convolution, the input data includes: the node feature matrix H of the first layer input l And the dynamic adjacency matrix, when the convolution is performed for the first time, the initial feature matrix H0 Convolve with the current dynamic adjacency matrix, the initial feature matrix H 0 It is a matrix composed of original features. The original features are the features obtained by extracting the running data for the first time. For example, when the running data includes log data, indicator data and performance data, the initial feature matrix includes 3 rows, each row corresponds to a feature extraction result, including: log data feature extraction result X log , feature extraction results of indicator data X metric And the feature extraction results of traffic data X flow There are three types of convolution operations: the initial feature matrix consists of N columns, each corresponding to a type of cloud network device, where N is the total number of cloud network devices in the cloud network environment. During subsequent convolution operations, the convolution objects are the updated dynamic adjacency matrix and the output of the previous convolution. The process of dynamically updating the adjacency matrix can be described by the following formula: Where, β represents the traffic normalization factor (the historical traffic mean is used in the embodiment of this application), which is used to control the weight amplification ratio. ij (t) represents the traffic between cloud network device i and cloud network device j. If the traffic between cloud network device i and cloud network device j increases suddenly, then Traffic ij (t)>>β, weight A ′ ij (t) exponential amplification, strengthening the detection of abnormal propagation paths. In order to ensure the stability of weights, the method provided in the embodiment of the present application performs normalization operation on the dynamically updated adjacency matrix by row. The normalization operation process can be performed using the formula To express, where is the result of the normalization operation, N is the number of nodes (i.e. the number of cloud network devices); A ′ ik (t)(k=1,……,N) is the weight of each edge associated with node i.

[0048] In this embodiment, convolution is used to aggregate neighbor features and the historical state of the network to generate high-order features. The specific process is as follows: First, according to the formula Perform neighborhood aggregation, where is the normalized adjacency matrix, H (l) It is the feature matrix obtained by extracting features from the input data in the lth neural network layer of DMF-GNN, and its dimension is (N×d), where N is the number of cloud network devices, that is, the number of nodes in the dynamic association graph, and d is the length of each feature vector; the input data of the lth neural network layer is the output data of the (l-1)th neural network layer, and the input data of the first neural layer of DMF-GNN is the initial feature matrix H0 , W (l) It is a matrix generated according to the model parameters of DMF-GNN, used to change Next, the activation function is used to perform nonlinear transformation on B1. The process of nonlinear transformation can be described as the formula In this formula, B2 is the neighborhood aggregation result (B1) after nonlinear transformation. In this embodiment, the activation function (ReLU) is used to perform nonlinear transformation on the neighborhood aggregation result (B1). x represents the object on which the nonlinear transformation is performed. In this embodiment, the object on which the nonlinear transformation is performed is the neighborhood aggregation result (B1). Furthermore, residual connection is used to retain the node's own historical state, avoid gradient disappearance, and enhance the stability of the model. By stacking layer by layer and gradually fusing multi-hop neighbor features, the node feature matrix H of the lth layer is l It follows the formula H (l+1) =B2+H (l) W (l) , the l in the formula can refer to any neural network layer in DMF-GNN. Finally, the high-order graph feature vector H is generated GCN , when the number of neural network layers contained in DMF-GNN is L, H GCN =H (L) , where H GCN Each row of H represents a cloud network device, characterizing its contextual state in the graph structure. If the feature vector of a row deviates significantly from the normal distribution, it is determined to be an abnormal device. GCN The column direction represents the high-order feature dimension, encoding spatiotemporal dependencies.

[0049] In step S204, for each device to be detected, the multimodal feature fusion vector is processed and analyzed using the abnormal probability analysis model to obtain a probability matrix output by the abnormal probability analysis model, wherein each element in the probability matrix is ​​used to indicate the probability of a type of abnormality existing in the device to be detected. The abnormal probability analysis model is an initial neural network model in which multiple neural network layers are deleted during the training process.

[0050] The method provided in the embodiment of the present application uses a neural network model (i.e., anomaly probability analysis model) trained by dynamic pruning technology to process the multimodal feature fusion vector generated based on multiple types of operating data, avoiding the problem of the inability to process multimodal heterogeneous data in the related art. When the dynamic pruning method is used for model training, unimportant neural network layers in the initial neural network model used as the training object are deleted during the training process. In step S204, the anomaly probability analysis model outputs a probability matrix after processing and analyzing the multimodal feature fusion vector. This probability matrix is ​​used to indicate the probability analysis of the anomaly type of the device to be detected, that is, each element in the probability matrix is ​​used to indicate the probability of a class of anomalies occurring in the device to be detected. The row elements of the probability matrix are used to indicate the device to be detected, and the column elements are used to indicate the probability of the occurrence of the anomaly type. If there is only one device to be detected, the probability matrix is ​​a column matrix, and each row of the column matrix corresponds to one anomaly type; if there are multiple devices to be detected, each column of the probability matrix corresponds to one device to be detected, and each row corresponds to one anomaly type.

[0051] In step S204, the abnormality probability analysis model can be loaded into the memory. For example, the raw data of the abnormality probability analysis model can be loaded from the non-volatile memory into the volatile memory so that the processor can run the abnormality probability analysis model. The raw data of the abnormality probability analysis model refers to unprocessed data, which generally includes parameters and structural data of the abnormality probability analysis model. The structural data can be a calculation relationship based on the parameters, such as the forward propagation calculation relationship between intermediate layers and neurons. Specifically, the structural data can include code related to the structure of the abnormality probability analysis model, such as code for performing related calculations between intermediate layers and neurons.

[0052] In one embodiment, a memory area for loading the abnormality probability analysis model can be divided, including a structure data storage area and a parameter storage area. The structure data storage area is used to store structure-related code, and the parameters referenced by it can point to the addresses of specific parameters in the parameter storage area through pointers. During the training process of the abnormality probability analysis model, the parameters may need to be frequently updated, and the parameter values ​​in the parameter storage area can be simply updated.

[0053] According to some optional embodiments of the present application, an abnormal probability analysis model is used to process and analyze the multimodal feature fusion vector to obtain a probability matrix output by the abnormal probability analysis model, including: determining multiple position coding vectors, and fusing the multiple position coding vectors and multiple multimodal feature fusion vectors to obtain a feature sequence, wherein the feature sequence is composed of multiple time series feature vectors, each time series feature vector is generated based on a position coding vector and a multimodal feature vector, each position coding vector is used to indicate the relative position of the corresponding time series feature vector in the feature sequence, and the position coding vector is generated during the training process of the abnormal probability analysis model; determining a processing order based on multiple relative positions, and processing each time series feature vector in the feature sequence in turn according to the processing order in the probability prediction layer until all the time series feature vectors are processed to obtain a hidden layer feature vector, wherein when processing the next time series feature vector, the processing result of the previous time series feature vector and the next time series feature vector are used together as the probability prediction layer input data; performing dimension mapping and normalization on the hidden layer feature vector to obtain a probability matrix.

[0054] Figure 3 It is a schematic diagram of the data processing process of the abnormal probability analysis model. In this embodiment, the multimodal feature fusion vector is used as the input data of the abnormal probability analysis model. Figure 3 As shown, in this embodiment, the process of processing and analyzing the multimodal feature fusion vector by the trained abnormal probability analysis model includes the following: First, the input multimodal feature vector and the position coding vector determined during the training process are fused at the position coding layer, wherein the position coding vector helps the abnormal probability analysis model understand the sequence relationship of different device features in the time series. The position coding vector can be automatically generated during the abnormal probability analysis model training process, and can be randomly initialized or generated according to a predefined function (such as sine and cosine functions) at the beginning. For example, sine and cosine functions can be used to encode position information to increase the model's sensitivity to position. In this embodiment, the fusion processing process can be performed using formula X temp =F input +P, where P represents the matrix composed of position encoding vectors, F input Represents a matrix composed of multiple multimodal feature vectors, X temp It is the result of the fusion process (i.e., feature sequence). Each element in the feature sequence is a time series feature vector generated by the fusion of a multimodal feature vector and a position encoding vector. Multiple time series feature vectors are arranged into a feature sequence according to the relative position defined by the position encoding vector. Therefore, the position encoding vector in each time series feature vector can indicate its relative position in the feature sequence. Figure 3As shown in the figure, after the position encoding layer completes the fusion process, the probability prediction layer (simplified LSTM) processes the feature sequence output by the position encoding layer. In this embodiment, before the probability prediction layer processes the feature sequence, the time series feature vector obtained by the fusion process is mapped to a unified dimension. The process of unifying the dimensions can be described by the formula: X1 = W p Flatten(X temp )+b p , in the formula, W p Represents the weight matrix of the fully connected layer, which is the model parameter of the abnormal probability analysis model, b p Represents the bias vector, which is the model parameter of the abnormal probability analysis model, Flatten represents the feature flattening operation, and X1 is the processing result, which is the X with changed dimension. temp Through the same dimension mentioned above, the static features are converted into pseudo time series signals X1 with time series features to adapt to the time series processing capabilities of the probability prediction layer. Figure 3 As shown in FIG, the probability prediction layer (simplified LSTM) is a LSTM network that is pruned from a standardized long short-term memory network. Therefore, in this embodiment, the abnormal probability analysis model including the simplified LSTM is also called a lightweight long short-term memory network (TinyLSTM) model. When the probability prediction layer performs data processing, the input data is the feature sequence X1 after dimension conversion, and the output hidden layer feature vector is h LSTM =LSTM(X1), LSTM(X1) is Figure 3 The complete process of the pruned LSTM processing the input X1 is shown. Since the gating mechanism (input gate i t 、Forget Gate t and output gate O t ) accesses the cell state using a peephole connection to achieve parameter compression and improve processing speed. Figure 3 The input gate i in the simplified LSTM t 、Forget Gate t , output gate O t The calculation no longer depends directly on the cell state (C t ), the specific data processing process represented by LSTM(X1) can be expressed by the following formula: Where W f is the weight matrix of the forget gate, b f is the bias vector of the forget gate; W i is the weight matrix of the input gate, b i is the bias vector of the input gate; W gis the weight matrix of the candidate memory, which is the candidate information added to the unit state, b g is the bias vector of the candidate memory; W o is the weight matrix of the output gate, b o is the bias vector of the output gate; the above parameters are all model parameters of the abnormal probability analysis model and remain unchanged after the model training is completed. σ is a function used to limit the output value to [0,1]. t is the data input to the streamlined LSTM at the current time (t), h t-1 is the hidden state vector at the previous moment (t-1). The processing order followed by the streamlined LSTM when processing the input feature sequence is determined by the relative position defined by the position encoding vector. The processing order refers to the processing order of each feature vector in the feature sequence. The streamlined LSTM processes the feature vectors in the feature sequence according to the processing order to obtain the hidden layer feature vector. The hidden layer feature vector is the hidden layer feature vector h output after the last time series feature vector in the feature sequence is processed. LSTM .

[0055] Still Figure 3 As shown, after obtaining the hidden layer feature vector h LSTM After that, in the fully connected layer, h LSTM Perform dimension mapping and transform the hidden state h of LSTM LSTM Mapped to the category space, unnormalized probabilities (logits) are generated; the activation function (softmax function) is then used to normalize the logits to obtain a probability matrix for describing the probability distribution of abnormal types. In this embodiment, the fully connected layer is used to calculate the probability distribution of h LSTM When performing dimension mapping, the weight matrix and bias vector of the fully connected layer are used to process logits. The dimension mapping process can be expressed as follows: logist =W FC ·h LSTM +b FC To describe, where Z logist It is the result of dimension mapping of logits and is an unnormalized probability vector. FC is the weight matrix of the fully connected layer, b FC is the bias vector of the fully connected layer, W FC and b FC These are all model parameters of the abnormal probability analysis model. Next, the hidden layer feature vector after dimension mapping is normalized. Specifically, a preset function (high temperature softmax function) is used to normalize the hidden layer feature vector Z after dimension mapping. logistThe value of each element in is scaled to complete the normalization process. In this embodiment, the normalization process can be performed using the formula To express, where p (t) student is the probability matrix, T (t) It is provided by the softmax function for Z logist The temperature parameter by which the value of each element in T is scaled. (t) It is determined during the training process of the abnormal probability analysis model. In this embodiment, in order to dynamically balance knowledge transfer and task optimization, a dynamic temperature adjustment strategy is adopted to improve the model accuracy. start = 5, soften the probability distribution and let the student model (the initial neural network model used to train the abnormal probability analysis model) imitate the coarse-grained category relationship of the teacher model; in the later stage of training, set the low temperature T end = 1, sharpens the probability distribution and focuses on fine-grained classification boundary optimization. During the entire training process, the strategy for dynamically adjusting the temperature is as follows: Where, T start is the initial temperature used in the early stage of training, T end is the final temperature used in the later stages of training, N epochs It is the number of iterations during model training.

[0056] According to some other optional embodiments of the present application, the abnormal probability analysis model is trained by the following method: obtaining a training data set, wherein the training data set includes: historical operation data and multiple real labels, the historical operation data are multiple types of operation data generated by multiple IoT devices in the historical operation period, and the real labels are used to indicate the operation status of the IoT devices in the historical operation period, and the operation status includes: normal operation and abnormal operation; performing feature fusion processing on the historical operation data to obtain a historical multimodal feature fusion vector; inputting the historical multimodal feature fusion vector into the teacher model and the initial neural network model respectively to obtain the prediction result output by the initial neural network model and the soft label output by the teacher model, wherein the teacher model is a neural network model with a larger number of parameters than the initial neural network model, the soft label is the operation status of the IoT device in the historical operation period predicted by the teacher model, and the prediction result is the operation status of the IoT device in the historical operation period predicted by the initial neural network model; determining a first loss function based on the soft label and the prediction result, determining a second loss function based on the prediction result and the real label, and determining a total loss function based on the first loss function and the second loss function; determining the training process of the initial neural network model based on the total loss function, wherein the training process includes: continuing iterative training and stopping iterative training.

[0057] In the solution provided in the embodiment of the present application, an adaptive lightweight edge inference framework (ALEIF) is introduced to process multimodal feature vectors, and finally output the abnormal probability distribution of the device to realize abnormal detection of cloud network devices. Under ALEIF, the data processing process includes the steps of feature standardization, knowledge distillation, dynamic pruning, and edge reasoning, wherein knowledge distillation and dynamic pruning occur during the training process of the abnormal probability analysis model. Among them, knowledge distillation refers to the process in which the student model learns the output distribution of the teacher model during the model training process, and transfers the knowledge of the teacher model to the student model, thereby realizing model compression and acceleration. In this embodiment, the student model refers to the initial neural network model that becomes the abnormal probability analysis model after training, and the teacher model is a neural network model with more parameters and deeper learning depth. In order to realize knowledge distillation, during the training process of the abnormal probability analysis model, both the teacher model and the student model process the training data, and the model training can be performed on the cloud server. In this embodiment, the set of training data (i.e., training data set) includes: various operating data generated by various Internet of Things (IoT) devices during historical operating periods, and real labels that record the operating status of the IoT devices during these historical operating periods. In this embodiment, the multiple real labels in the training data set include: a class of real labels indicating normal operating status and a class of real labels indicating abnormal operating status. During the training process, the historical operating data serves as input data for the student model and the teacher model. The student model processes the input data and outputs a prediction result, which is the operating status of the IoT device during the historical operating period predicted by the student model based on the IoT's historical operating data. The prediction result output by the teacher model after processing the input data is called a soft label, which records the operating status of the IoT device during the historical operating period predicted by the teacher model based on the IoT's historical operating data. Next, the student model is trained using the real labels in the training data set and the soft labels output by the teacher model. The student model learns the coarse-grained category classification of the teacher model and transfers the knowledge in the teacher model (i.e., the method for predicting a certain category of anomaly for each device to be tested) to the student model, achieving knowledge distillation. Specifically, a mixed loss function (i.e., total loss function) is determined based on the true label, soft label, and prediction results output by the initial neural network model. The value of the mixed loss function determines whether to end or continue iterative training. The loss function calculated based on the soft label and prediction results (i.e., the first loss function) is The first loss function L KL It is used to minimize the difference between the prediction results and the soft labels, promote the learning of the student model, and make the probability distribution of its output as close as possible to the output of the teacher model, thereby realizing the knowledge transfer of the teacher model;(t) is the temperature parameter, and Figure 3 The temperature parameters used in the normalization process are the same; teacher It is a soft tag, P student (t) Is the prediction result output by the student model during the tth iteration of training. The cross entropy loss function (i.e., the second loss function) is determined based on the prediction result and the true label. The second loss function L CE Used to minimize the difference between the predicted results and the true labels, making the predicted results as close to the true labels as possible; is the true label of the i-th category, i is the type of abnormal category, and each value of i represents a category. In this embodiment, the number of abnormal categories recorded in the true label is 9, so the maximum value of i is 9. Hybrid loss function (i.e., total loss function) L total =0.7·L KL +0.3·L CE ; in L total When the difference between the value of and the preset loss function value remains unchanged for several consecutive iterations, the iterative training is stopped, otherwise, the iterative training is continued. According to the method provided in this embodiment, the knowledge of the complex neural network model (i.e., the teacher model) is transferred to the lightweight student model through knowledge distillation, which effectively solves the problems of high inference delay and large memory usage of the complex neural network model, and effectively avoids the performance bottleneck caused by insufficient capacity when the lightweight student model is directly trained. At the same time, the generalization ability and tolerance to noise of the lightweight student model are enhanced, achieving a balance between high efficiency and performance.

[0058] In this embodiment, a deep multi-layer perceptron model (Multi-Layer Perceptron, Deep MLP) is used as the teacher model. Figure 4 This is a schematic diagram of the model structure of the teacher model, such as Figure 4 As shown in Figure 2, the teacher model consists of four fully connected layers, each of which is connected in series with different activation functions and regularization techniques. The data of the training set is used as the input data F of the Deep MLP model. inputThe specific processing of each layer is as follows. The input data is sent to the first layer structure (FC1). The dimension of the fully connected layer in the first layer structure is 512 dimensions. The activation function used is the ReLU function, and the regularization technology used is Dropout (p = 0.3). The first layer structure is used to perform nonlinear expansion of the input features to prevent overfitting. The 256-dimensional input features are linearly mapped to the 512-dimensional space through the fully connected layer (512 neurons) to enhance the feature expression ability; the ReLU activation function is used to introduce nonlinearity to solve the problem of insufficient expression ability of the linear model, filter negative values, enhance sparsity, and alleviate gradient disappearance; the Dropout layer is set to p = 0.3, and 30% of neurons are randomly discarded to prevent overfitting; the data processing process of the first layer is described by the following formula: Among them, FC1 is the output of the first layer; W1 is the weight matrix of the first layer, b1 is the bias vector of the first layer, W1 and b1 are model parameters, and ReLU is used to correct the value of the data x it processes to 0 or x. The second layer structure (FC2) is used to extract high-order features, such as Figure 4 As shown, the dimension of the fully connected layer in the second layer structure is 256 dimensions, the activation function used is the ReLU function, and the regularization technology used is batch normalization (BatchNorm). The output of the first layer is sent to the second layer, and the 512-dimensional features are compressed to 256 dimensions through the fully connected layer (256 neurons) to extract higher-order features; the ReLU activation function is used to maintain nonlinear expression capabilities; the batch normalization operation is used to achieve standardized output distribution, accelerate network convergence, and stabilize the training process. The data processing process of the second layer (FC2) can be described as the formula FC2 = BatchNorm (ReLU (W2FC1 + b2)), where FC2 is the output of the second layer; W2 is the weight matrix of the second layer, b2 is the bias vector of the second layer, and W2 and b2 are model parameters. The third layer structure (FC3) is used to further extract features and output low-dimensional features; as shown Figure 4 As shown in the figure, the dimension of the fully connected layer in the third layer structure is 128 dimensions, and the activation function used is the GELU activation function; the result of the second layer is input to the third layer, and the 256-dimensional features are compressed to 128 dimensions through the fully connected layer (128 neurons) to further extract features; the GELU activation function is used to alleviate the problem of "some neurons permanently stop learning" of ReLU, improve the robustness of the model, and capture complex patterns through the smooth non-characteristics of GELU, providing refined low-dimensional features for the output layer (FC4). The data processing process of the third layer (FC3) can be described as the formula Among them, FC3 is the output of the third layer; W3 is the weight matrix of the third layer, b3 is the bias vector of the third layer, and W3 and b3 are model parameters. X is the data processed by GELU, which is (ReLU(W3FC2+b3)) in this embodiment. φ(x) represents that GELU uses the cumulative distribution function of the standard normal distribution to process the data. The last layer (FC4) consists of a fully connected layer and a softmax high temperature function. The output FC3 of the third layer is sent to the last layer. The 128-dimensional features are mapped to a 9-dimensional category space through a fully connected layer (9 neurons). Softmax high temperature softening (T=3) is used to soften the probability distribution during the training phase to make the small probability category more significant and transfer the correlation between different abnormal categories. The data processing process of the last layer can be described as the formula Among them, q teacher is the output of the teacher model, FC3 is the output of the third layer, W4 is the weight matrix of the fourth layer, b4 is the bias vector of the fourth layer, W4 and b4 are model parameters; softmax T is the softmax high temperature function, where softmax T (x) i Represents the probability of the occurrence of the i-th abnormal category, and the x in the brackets is the softmax T In this embodiment, the softmax T The processed data is (W4FC3+b4), T is the temperature parameter used to smooth the output of the Softmax function; x i Represents the output of the i-th neuron, softmax T (x) i The probability of the existence of the i-th type of anomaly is represented by j, and when the values ​​of i and j are equal, i and j represent the same type of anomaly. The soft label q finally output by the teacher model (Deep MLP model) is teacher It is a 9-dimensional vector representing the probability distribution, which indicates the probability of the existence of anomalies in each category.

[0059] According to some optional embodiments of the present application, multiple neural network layers of the initial neural network model are deleted during the training process by the following method: obtaining the absolute value of the weight of each layer in the initial neural network model, and arranging the multiple absolute values ​​into a weight sequence in ascending order; determining a pruning threshold based on the weight sequence, wherein the pruning threshold is used to assist in determining the deleted neural network layer; assigning an absolute value less than the pruning threshold to an invalid value, and assigning an absolute value greater than the pruning threshold to a valid value; determining the neural network layer corresponding to the invalid value as the deleted neural network layer, until the sparsity of the initial neural network model reaches a preset sparsity, stopping the iteration, and obtaining an abnormal probability analysis model, wherein the preset sparsity is used to indicate the ratio of the number of deleted neural network layers to the number of all neural network layers contained in the initial neural network model.

[0060] To ensure that the training process reduces redundant parameters while maintaining model performance, this example uses a dynamic pruning method to determine which neurons in the initial neural network model are deleted during the training of the anomaly probability analysis model (TinyLSTM model). Given that a fixed threshold cannot adapt to changes in the distribution of model weights during training, an exponential moving average (EMA) threshold is introduced to automatically track changes in weight distribution. This smoothing threshold adjustment suppresses sudden changes in pruning intensity, ensuring the stability of model training. Figure 5 is the pseudo code of the dynamic pruning algorithm, such as Figure 5As shown, in each iteration, the batch size needs to be counted every pruning interval to obtain the absolute value of the weight of the prunable layer in the initial neural network model (here, each layer of the initial neural network model is considered to be a prunable layer); or, when it is pre-defined that only the convolution layer is the prunable layer, the absolute value of the convolution layer in the initial neural network model is obtained. After obtaining the absolute value of the weight of the prunable layer, the weights of all prunable layers are merged into a weight vector (i.e., a weight sequence). Specifically, the absolute values ​​of the weights are arranged in ascending order from small to large to obtain the weight vector. The EMA threshold (i.e., the pruning threshold) is dynamically updated based on the weight sequence. The EMA threshold is used to assist in determining the deleted neural network layer. Specifically, the absolute value of the weight less than the EMA threshold is marked as an invalid value, such as set to 0; the absolute value of the weight greater than or equal to the EMA threshold is marked as a valid value. After completing the above-mentioned weight validity evaluation, the weights corresponding to the invalid values ​​are actually set to zero in the model, which is equivalent to "deleting" the neuron connections associated with these invalid weights; repeat the above-mentioned pruning process until the sparsity of the initial neural network model reaches the preset sparsity, and the abnormal probability analysis model can be obtained; wherein, the sparsity of the initial neural network model refers to the ratio of the number of neural network layers set to 0 (i.e., deleted) in the initial neural network model to the total number of neural network layers contained in the initial neural network model. For example, when the preset sparsity is set to 70%, when the proportion of neural network layers set to 0 (i.e., deleted) in the initial neural network model reaches 70%, the iterative pruning is stopped and the abnormal probability analysis model is output.

[0061] Optionally, determining a pruning threshold based on the weight sequence includes: during the first iterative training, determining the deletion ratio corresponding to each absolute value, and when the deletion ratio is equal to the preset sparsity, determining the absolute value corresponding to the deletion ratio as the pruning threshold, wherein the deletion ratio is used to indicate the ratio of the number of other absolute values ​​in the weight sequence that are smaller than the absolute value to the total number of absolute values ​​contained in the weight sequence; during non-first iterative training, updating the weight sequence according to the results of the previous iterative training to obtain a new weight sequence; determining the deletion ratio corresponding to each absolute value in the new weight sequence, and when the deletion ratio is equal to the preset sparsity, determining the absolute value corresponding to the deletion ratio as the current pruning threshold; and determining the pruning threshold based on the historical pruning threshold and the current pruning threshold used in the previous iterative training process.

[0062] Considering that the fixed threshold cannot adapt to the distribution changes of model weights during training, the EMA threshold is introduced to automatically track the changes in weight distribution. Therefore, the EMA threshold used in the embodiment of the present application changes dynamically with the iterative training of the model. Specifically, Figure 5As recorded in the background, in the first iteration, the quantile of the weight sequence is determined as the EMA threshold (i.e. the pruning threshold); in subsequent iterations, the EMA threshold is updated according to the changes of the neural network model. Wherein, when the quantile of the weight sequence is determined as the EMA threshold, the absolute values less than the quantile in the weight sequence are determined, and the neural network layers corresponding to these absolute values less than the quantile are the neural network layers to be deleted. In each dynamic pruning, the total number of the deleted neural network layers and the neural network layers contained in the initial neural network model reaches the preset sparsity, therefore, when determining the quantile, the deletion ratio corresponding to each absolute value in the weight sequence is determined respectively until a deletion ratio identical to the preset sparsity is found, and the absolute value corresponding to this deletion ratio identical to the preset sparsity is determined as the quantile (i.e. the EMA threshold used in the first iteration); for example, if the preset sparsity is 0.7, the absolute value corresponding to the deletion ratio of 0.7 is determined as the quantile. The deletion ratio corresponding to each absolute value is determined by the following method: in the weight sequence, the number of other absolute values whose values are less than the absolute value is determined, and the ratio of the number of the other absolute values to the total number of absolute values contained in the weight sequence is determined as the deletion ratio corresponding to the absolute value. In subsequent iteration training (i.e. non-first iteration training), the EMA threshold is updated according to the formula θ Figure 5 t = β·θ t-1 +(1-β)θ current , wherein θ t represents the update result of the EMA threshold, β is the EMA smoothing coefficient which can be set according to actual needs, θ t-1 is the EMA threshold used in the last iteration training process (i.e. the historical pruning threshold), and θ current is the quantile (i.e. the current pruning threshold) determined in the weight sequence (i.e. the new weight sequence) according to the preset sparsity in the current round. Since part of the neural network layers in the initial neural network model are deleted after each round of iteration training, the weight sequence is updated after each round of iteration training. The method for determining the quantile in the new weight sequence is the same as the method for determining the quantile in the first iteration. In the non-first iteration training, the pruning threshold used is θ t .

[0063] In step S206, it is judged whether the to-be-detected device is abnormal in the preset time window according to the plurality of probability matrices corresponding to the preset time window, and a judgment result is obtained, wherein the preset time window contains a plurality of continuous detection periods, and the judgment result records the abnormal type of the to-be-detected device.

[0064] ​The method provided in the embodiment of the present application determines whether there is an abnormality in the device by comprehensively considering the operating status of the device to be detected within a period of time; then in step S206, for a detection time interval (i.e., a preset time window) composed of multiple continuous detection time periods, multiple probability matrices corresponding to the time window are obtained, wherein the number of detection time periods included in the preset time window is pre-set, the preset time window contains multiple detection time periods, and each detection time period contains multiple detection moments; each detection moment corresponds to a probability matrix, and therefore, each preset time window corresponds to multiple probability matrices. A probability matrix corresponding to each detection moment is obtained by processing the multimodal feature fusion vector generated by the multiple operating data generated by the device to be detected at the detection moment by the abnormal probability analysis model. For example, the preset time window can be set to 10 seconds, the detection period can be set to 10 seconds, and the detection moment can be set to 1 second. Then, the method provided in the embodiment of the present application is adopted, and each preset time window corresponds to 10 probability matrices. The multiple probability matrices corresponding to the preset time window are updated as the new probability matrix is ​​generated. Specifically, after obtaining multiple types of operating data of the device to be detected in real time and processing to obtain a new probability matrix, the new probability matrix is ​​added to the probability matrix corresponding to the preset time window, and the detection moment farthest from the current moment is removed from the preset time window, and the probability matrix corresponding to the detection moment farthest from the current moment is removed from the multiple probability matrices corresponding to the preset time window. Furthermore, based on the multiple probability matrices corresponding to the preset time window, it is judged whether the device to be detected has an operating anomaly within the duration corresponding to the preset time window, wherein the judgment result also records the specific anomaly type, such as network attack, performance bottleneck, hardware failure, etc.

[0065] Optionally, whether the device to be detected has an abnormality is judged based on multiple probability matrices corresponding to a preset time window to obtain a judgment result, including: for each probability matrix, determining a target element in the probability matrix that is greater than a preset probability value, and when the target element exists in any probability matrix, determining the judgment result as that the device to be detected has an abnormality, and determining the abnormality type corresponding to the target element as the abnormality type of the device to be detected, and the abnormality types include: network attack, performance bottleneck, hardware failure; when the target element does not exist in multiple probability matrices corresponding to the preset time window, determining the judgment result as that the device to be detected has no abnormality.

[0066] In view of the fact that single abnormality probability triggering alarm is easily affected by transient noise (such as short-time flow fluctuation, sensor false alarm) and thus leads to high false alarm rate, the method provided in the embodiments of the present application adopts a multi-window cumulative alarm method to balance sensitivity and false alarm rate by using a sliding time window and a cumulative triggering technique, thereby improving the robustness of the alarm. In the embodiments, the elements contained in each probability matrix are used to indicate the probability that a certain abnormality exists in a certain to-be-detected device, and a triggering condition (i.e., a preset probability value) of a single abnormality probability in each preset time window is set in advance, for example, set to 0.8. The method for determining whether the to-be-detected device is abnormal according to the plurality of probability matrices is as follows: if there is a probability value greater than 0.8 in any one of the probability matrices, the probability value greater than 0.8 is determined as a target element, it is determined that the to-be-detected device is abnormal when running in the time period corresponding to the preset time window, and the abnormality type is the abnormality type corresponding to the probability value greater than 0.8. If there is no probability value greater than the preset probability value (i.e., the target element) in the plurality of probability matrices corresponding to the preset time window, it is indicated that the to-be-detected device is not abnormal when running in the time period corresponding to the preset time window; however, the to-be-detected device may be abnormal after the preset time window is updated, and in the embodiments, after the preset time window is updated each time, whether the target element greater than the preset probability value exists in the probability matrix corresponding to the detection time period newly added in the preset time window is detected to re-detect whether the to-be-detected device is abnormal in the detection time period corresponding to the updated preset time window. The sliding window is dynamically updated: each time a new detection time period is added to the window, the detection time period farthest from the current time is removed from the window, and the number of triggering times is dynamically updated, thereby ensuring the real-time performance of the alarm response.

[0067] In step S208, the alarm levels of each abnormality type are determined according to the plurality of determination results corresponding to the plurality of preset time windows.

[0068] The method provided in the embodiments of the present application adopts the multi-window cumulative alarm method to determine the final output alarm level, and thus in step S208, the determination results corresponding to each preset time window in the continuous plurality of preset time windows are obtained, and the alarm levels of each abnormality type are determined according to the obtained plurality of determination results; the above method effectively reduces the false alarm rate.

[0069] Optionally, the alarm level of each abnormality type is determined based on multiple judgment results corresponding to multiple preset time windows: the number of alarms corresponding to each abnormality type is determined, wherein the number of alarms is the number of times the abnormality type is recorded in multiple judgment results; the alarm level of the abnormality type corresponding to the number of alarms greater than or equal to the first preset value is determined as the first level; the alarm level of the abnormality type corresponding to the number of alarms greater than or equal to the second preset value is determined as the second level, wherein the second preset value is greater than the first preset value, and the second level is higher than the first level; the alarm level of the abnormality type corresponding to the number of alarms greater than or equal to the third preset value is determined as the third level, wherein the third preset value is greater than the second preset value, and the third level is higher than the second level.

[0070] The method provided in the embodiment of the present application is to confirm the alarm only when the cumulative number exceeds the threshold value by monitoring the number of triggers of the abnormal probability (i.e., the probability greater than the preset probability value) in the continuous time window, and effectively filter out sporadic noise. In this embodiment, when monitoring the number of triggers of the abnormal probability in the continuous time window, for each preset time window in the continuous time window, determine the abnormal type recorded in the judgment result of the preset time window, which is the type of abnormality indicated by the abnormal probability; further, determine the number of times each abnormal type appears in the multiple judgment results corresponding to the multiple continuous time windows, and record it as the alarm number corresponding to the abnormal type, so as to determine the alarm level corresponding to the abnormal type according to the alarm number. Specifically, in this embodiment, the corresponding alarm number is set for different alarm levels, wherein the higher the alarm level, the more alarm number it corresponds to. For example, the number of preset time windows can be set to 10, the length of each preset time window is 2 seconds, the alarm number corresponding to the first level (low-level alarm) is set to 2 (i.e., the first preset value), the alarm number corresponding to the second level (intermediate alarm) is set to 3, and the alarm number corresponding to the third level (high-level alarm) is set to 4. When within a preset time window, the probability value corresponding to a certain abnormal type exceeds the preset probability value, and the cumulative number is 1; based on the above, the number of times the probability value of each abnormal type exceeds the preset probability value within multiple preset time windows (i.e., the number of alarms); if the number of alarms for a certain abnormal type is greater than or equal to 2 (i.e., the first preset value) and less than 3 (the second preset value), a first-level alarm (low-risk alarm) is issued for this abnormal category; if the number of alarms for a certain abnormal type is greater than or equal to 3 (i.e., the second preset value) and less than 4 (the third preset value), a second-level alarm (medium-risk alarm) is issued for this abnormal category; if the number of alarms for a certain abnormal type is greater than or equal to 4 (i.e., the third preset value), a third-level alarm (high-risk alarm) is issued for this abnormal category. The alarm method can be to send text messages to the on-duty personnel and administrators.

[0071] It should also be noted that in the method provided in the embodiment of the present application, an alarm threshold can also be set in advance. The alarm threshold is a probability value greater than the preset probability value. For example, when the preset probability value is 0.8, the alarm threshold can be set to 0.95. When any probability value in the multiple probability matrices corresponding to the preset time window is greater than the alarm threshold, the alarm information is directly output without the need for multi-window accumulation.

[0072] Through the above steps, it is possible to effectively integrate different multimodal heterogeneous data (i.e., multiple types of operating data), and adapt to the dynamic changes of the cloud network environment; a lightweight anomaly probability analysis model is used to process the fused multimodal heterogeneous data (i.e., multimodal feature fusion vector), predict the abnormal operation of the equipment to be detected and the abnormal type of the equipment to be detected, thereby improving the accuracy of the prediction results; the anomaly frequency is counted through a sliding window, and the alarm threshold is dynamically adjusted in combination with the output probability distribution, supporting multi-level alarm and response disposal, improving the robustness of the alarm and reducing the false alarm rate.

[0073] Figure 6 This is a structural diagram of the device for detecting cloud network equipment provided by the embodiment of the present application, such as Figure 6 As shown, the device for detecting cloud network devices includes: an acquisition module 60, which is used to obtain multiple operating data of the device to be detected during the detection period, and generate a multimodal feature fusion vector based on the multiple operating data, wherein the device to be detected is a cloud network device in a cloud network environment, and the number of devices to be detected is multiple; a processing module 62, which is used to use an abnormality probability analysis model to process and analyze the multimodal feature fusion vector for each device to be detected, and obtain a probability matrix output by the abnormality probability analysis model, wherein each element in the probability matrix is ​​used to indicate the probability of a class of abnormality existing in the device to be detected, and the abnormality probability analysis model is an initial neural network model with multiple neural network layers deleted during the training process; a judgment module 64, which is used to judge whether the device to be detected has an abnormality in the preset time window based on multiple probability matrices corresponding to the preset time window, and obtain a judgment result, wherein the preset time window contains multiple continuous detection periods, and the judgment result records the abnormality type of the device to be detected; a determination module 66, which is used to determine the alarm level of each abnormality type based on multiple judgment results corresponding to multiple preset time windows.

[0074] Figure 7 This is a workflow diagram of the device for detecting cloud network equipment, such as Figure 7As shown, when the device for detecting cloud network equipment is working, the acquisition module 60 collects various types of operating data generated by the cloud network equipment during the detection period, such as log data, performance indicator data, network traffic data, etc. The acquisition module 60 further performs fusion processing on the collected operating data to obtain a multimodal feature fusion vector; wherein, the process of fusion processing to generate a multimodal feature fusion vector includes data preprocessing, construction of a spatiotemporal semantic association graph (i.e., a dynamic association graph), and encoding of the spatiotemporal dependency characteristics of device behavior; data preprocessing is to convert various types of operating data into structured data; when constructing the spatiotemporal semantic association graph, the cloud network equipment is used as a node, and the edges connecting the nodes are constructed according to the spatiotemporal association relationship and semantic association relationship between the cloud network devices; the spatiotemporal dependency characteristics between the cloud network devices are extracted by encoding in the dynamic association graph, and finally a multimodal feature fusion vector is generated. The processing module 62 is used to receive the multimodal fusion feature vector, use the TinyLSTM model (i.e., the abnormal probability analysis model) that has undergone knowledge distillation and dynamic pruning optimization to process and analyze the multimodal fusion feature vector, and output the probability distribution (i.e., probability matrix) of various types of abnormalities (such as network attacks, performance bottlenecks, hardware failures) that may exist in the device to be detected during the current detection period. The judgment module 64 determines whether there is a target element (i.e., an abnormal type with a probability value greater than a preset probability value) based on multiple probability matrices corresponding to a preset time window (so the length of the time window can be 10 seconds, 30 seconds, 1 minute, etc.). If the target element is detected in any time window, one alarm of the abnormal type is recorded; if the target element is not detected in all time windows, it is considered that the device is operating normally. The determination module 66 dynamically adjusts its alarm level based on the number of alarms for each abnormal type in multiple preset time windows. For example: monitor 10 consecutive preset time windows. If the duration of each preset window is 1 second, and the number of alarms of a certain abnormal type within the 10-second window is greater than or equal to 2, then the alarm level is determined to be the first level (low-risk alarm); if the number of alarms of a certain abnormal type within the 10-second window is greater than or equal to 3, then the alarm level is determined to be the second level (medium-risk alarm); if the number of alarms of a certain abnormal type within the 10-second window is greater than or equal to 4, then the alarm level is determined to be the third level (high-risk alarm). Figure 7 As shown, after determining the alarm level, the device for detecting cloud network equipment can also automatically or semi-automatically trigger the corresponding response and disposal process according to the determined alarm level. For example, low-level alarms can increase the monitoring frequency through automation tools, medium-level alarms notify operation and maintenance personnel to intervene in analysis, and high-level alarms initiate emergency shutdown or isolation operations to minimize the impact of abnormalities in the detection equipment.

[0075] It should be noted that Figure 6 The preferred implementation of the embodiment shown can be found in Figure 2 The relevant description of the illustrated embodiment will not be repeated here.

[0076] Figure 8 This is a schematic diagram of the execution framework of the method for detecting cloud network devices, such as Figure 8 As shown, the method for detecting cloud network devices provided in the embodiment of the present application can be implemented using a framework composed of an edge server, a cloud server and a terminal, wherein the edge server is used to implement the following method: obtain multi-source heterogeneous data of the device to be detected from the cloud network environment, including but not limited to log data, performance indicator data and network traffic data. These are preprocessed and converted into structured data. Call dynamic graph construction and multimodal feature fusion to map the preprocessed heterogeneous data into a dynamic graph neural network (DMF-GNN) to construct a spatiotemporal-semantic association graph (i.e., a dynamic association graph). Call the dynamic graph convolution module to extract high-order features of nodes in the dynamic association graph through graph convolution operations. Call the multimodal fusion module to encode the spatiotemporal dependency features of device behavior, and then generate a multimodal fusion feature vector. As shown Figure 8 As shown, the multimodal fusion feature vector processed by the DMF-GNN model is then subjected to feature vector standardization, and a standardized feature fusion vector (standardized multimodal feature vector) is generated by performing dimensional conversion and standardization. The edge server will call the trained student model (TinyLSTM) to process the standardized feature fusion vector and output the abnormal probability distribution (i.e., probability matrix). Furthermore, the edge server continuously monitors the output results of the standardized fusion feature vector and counts the number of abnormal probability triggers within the sliding time window. Based on the cumulative number of abnormalities and the probability distribution, the alarm threshold and level are dynamically adjusted to ensure that the operation and maintenance personnel are notified of the abnormal equipment through the terminal in a timely and accurate manner so that appropriate disposal measures can be taken. The above-mentioned trained student model is trained on the cloud server through knowledge distillation and dynamic pruning methods, such as Figure 8 As shown in the figure, on a cloud server, by migrating knowledge from a complex model (i.e., a teacher model, such as Deep MLP) to a lightweight model (i.e., a student model, such as TinyLSTM), model compression and acceleration are achieved while maintaining high detection performance. During training, the TinyLSTM model calculates a loss function based on the soft labels output by the teacher model, the predictions output by the student model, and the true labels in the training dataset. The student's parameters are continuously updated during training to improve the accuracy of its predictions. Dynamic pruning automatically prunes weights based on their importance. By generating pruning masks that simulate the effect of deleting neural network layers, it reduces redundant parameters and enhances the efficiency of model deployment on edge devices.

[0077] Figure 8In the process shown, when the multimodal fusion feature vector output by the DMF-GNN model is used as the input of the ALEIF architecture, in order to adapt the high-dimensional features (512 dimensions) of the multimodal fusion feature vector to the input dimension (256 dimensions) of the knowledge distillation model and ensure the consistency of data distribution, the multimodal fusion feature vector is uniformly processed in the feature standardization module. First, the multimodal feature fusion vector is reduced in dimension using the adaptation layer of the fully connected layer. The dimensionality reduction process can be described as formula F adapted =ReLU(W adapted F fusion +b adapted ), where F adapted It is the multimodal feature fusion vector after dimensionality reduction. ReLU is a nonlinear activation function that does not change the dimension of the feature vector. adapted is the weight matrix of the adaptation layer, which is used to map 512 features to 256 dimensions, F fusion represents the multimodal fusion feature vector, b adapted is the bias term of the adaptation layer. Secondly, the normalization operation is used to align the training data distribution to ensure that the distribution of the multimodal feature fusion vector is consistent with the training data distribution of the TinyLSTM model, eliminating the difference in feature dimensions, which is conducive to accelerating the convergence of the model; the normalization operation process can be achieved using the formula Where μ train is the average value of the historical running data in the training set data (calculated after conversion to structured data), σ train is the standard deviation of the historical running data in the training set data (calculated after conversion to structured data), F adapted It is the multimodal feature fusion vector after dimensionality reduction.

[0078] An embodiment of the present application also provides a non-volatile storage medium, in which a computer program is stored. The device where the non-volatile storage medium is located executes the above method for detecting cloud network devices by running the computer program.

[0079] The above-mentioned non-volatile storage medium is used to store a program that performs the following functions: obtaining multiple operating data of the device to be detected during the detection period, and generating a multimodal feature fusion vector based on the multiple operating data, wherein the device to be detected is a cloud network device in a cloud network environment, and the number of devices to be detected is multiple; for each device to be detected, using an abnormality probability analysis model to process and analyze the multimodal feature fusion vector to obtain a probability matrix output by the abnormality probability analysis model, wherein each element in the probability matrix is ​​used to indicate the probability of a type of abnormality existing in the device to be detected, and the abnormality probability analysis model is an initial neural network model with multiple neural network layers deleted during the training process; judging whether the device to be detected has an abnormality in the preset time window based on multiple probability matrices corresponding to the preset time window, and obtaining a judgment result, wherein the preset time window contains multiple continuous detection periods, and the judgment result records the abnormality type of the device to be detected; determining the alarm level of each abnormality type based on multiple judgment results corresponding to multiple preset time windows.

[0080] An embodiment of the present application also provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute the above method for detecting cloud network devices through the computer program.

[0081] The processor in the above-mentioned electronic device is used to run a program that performs the following functions: obtaining multiple operating data of the device to be detected during the detection period, and generating a multimodal feature fusion vector based on the multiple operating data, wherein the device to be detected is a cloud network device in a cloud network environment, and the number of devices to be detected is multiple; for each device to be detected, using an abnormality probability analysis model to process and analyze the multimodal feature fusion vector to obtain a probability matrix output by the abnormality probability analysis model, wherein each element in the probability matrix is ​​used to indicate the probability of a type of abnormality existing in the device to be detected, and the abnormality probability analysis model is an initial neural network model with multiple neural network layers deleted during the training process; judging whether the device to be detected has an abnormality in the preset time window based on multiple probability matrices corresponding to the preset time window, and obtaining a judgment result, wherein the preset time window contains multiple continuous detection periods, and the judgment result records the abnormality type of the device to be detected; determining the alarm level of each abnormality type based on multiple judgment results corresponding to multiple preset time windows.

[0082] An embodiment of the present application also provides a computer program product, including computer instructions, which, when executed by a processor, implement the steps of the above method for detecting cloud network devices.

[0083] It should be noted that the various modules in the above-mentioned device for detecting cloud network equipment can be program modules (for example, a set of program instructions that implement a certain specific function) or hardware modules. For the latter, it can be expressed in the following forms, but is not limited to this: the expression form of each of the above-mentioned modules is a processor, or the functions of each of the above-mentioned modules are implemented by a processor.

[0084] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0085] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0086] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0087] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0088] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0089] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the relevant technology or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0090] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A method for detecting cloud network equipment, characterized in that: include: Acquiring multiple operating data of a device to be detected during a detection period, and generating a multimodal feature fusion vector based on the multiple operating data, wherein the device to be detected is a cloud network device in a cloud network environment, and the number of the device to be detected is multiple; For each of the devices to be detected, the multimodal feature fusion vector is processed and analyzed using an abnormality probability analysis model to obtain a probability matrix output by the abnormality probability analysis model, wherein each element in the probability matrix is ​​used to indicate the probability of a type of abnormality existing in the device to be detected, and the abnormality probability analysis model is an initial neural network model with multiple neural network layers deleted during the training process; Determining whether the device to be detected has an abnormality in the preset time window according to the plurality of probability matrices corresponding to the preset time window, and obtaining a determination result, wherein the preset time window includes a plurality of consecutive detection time periods, and the determination result records the abnormality type of the device to be detected; The alarm level of each abnormality type is determined according to the multiple judgment results corresponding to the multiple preset time windows.

2. The method according to claim 1, characterized in that Generating a multimodal feature fusion vector based on multiple types of the operating data, including: Generating a dynamic association graph based on the multiple types of operating data, wherein the types of operating data include log data, indicator data, and traffic data, each node in the dynamic association graph represents a type of device to be detected, and the edges of the dynamic association graph are constructed based on the spatiotemporal association relationships and semantic association relationships between the devices to be detected; Performing feature extraction on the dynamic association graph to obtain a high-order feature vector of the cloud network device, wherein the high-order feature vector is used to describe multi-dimensional information of the plurality of devices to be detected, including: operating information of the devices to be detected and association relationships between the plurality of devices to be detected; The multimodal feature fusion vector is generated according to the high-order feature vector and the feature vector obtained by extracting features from each type of the operating data.

3. The method according to claim 2, characterized in that Feature extraction is performed on the dynamic association graph to obtain high-order feature vectors of cloud network devices, including: For each edge in the dynamic association graph, determine the two nodes associated with the edge, and obtain two sets of operating data corresponding to the two nodes; determine the semantic similarity of events occurring between the two devices to be detected corresponding to the two nodes within the detection period based on the two sets of operating data; determine the event whose semantic similarity is greater than a preset semantic similarity as a target event, and determine the target time at which the target event occurs between the two devices to be detected corresponding to the two nodes within the detection period based on the two sets of operating data; determine the spatiotemporal association strength of the two devices to be detected corresponding to the two nodes based on the target time; and determine the weight of the edge based on the spatiotemporal association strength and the semantic similarity; An adjacency matrix is ​​determined according to the multiple weights corresponding to the multiple edges; and a convolution process is performed on the adjacency matrix to obtain the high-order eigenvector.

4. The method according to claim 1, wherein The multimodal feature fusion vector is processed and analyzed using an abnormal probability analysis model to obtain a probability matrix output by the abnormal probability analysis model, including: Determine a plurality of position coding vectors, and fuse the plurality of position coding vectors with the plurality of multimodal feature fusion vectors to obtain a feature sequence, wherein the feature sequence is composed of a plurality of time series feature vectors, each of the time series feature vectors is generated based on a position coding vector and a multimodal feature vector, each of the position coding vectors is used to indicate the relative position of the corresponding time series feature vector in the feature sequence, and the position coding vector is used to be generated during the training process of the abnormality probability analysis model; Determining a processing order according to the plurality of relative positions, and processing each of the time series feature vectors in the feature sequence in sequence according to the processing order at the probability prediction layer until all the time series feature vectors are processed, thereby obtaining a hidden layer feature vector, wherein when processing the next time series feature vector, the processing result of the previous time series feature vector and the next time series feature vector are used together as input data for the probability prediction layer; Perform dimension mapping and normalization processing on the hidden layer feature vector to obtain the probability matrix.

5. The method according to claim 1, wherein The abnormal probability analysis model is trained by the following method: Obtaining a training data set, wherein the training data set includes: historical operation data and multiple real labels, the historical operation data being multiple types of operation data generated by multiple IoT devices during historical operation periods, and the real labels being used to indicate the operation status of the IoT devices during the historical operation periods, wherein the operation status includes: normal operation and abnormal operation; Performing feature fusion processing on the historical operation data to obtain a historical multimodal feature fusion vector; The historical multimodal feature fusion is input into the teacher model and the initial neural network model respectively to obtain a prediction result output by the initial neural network model and a soft label output by the teacher model, wherein the teacher model is a neural network model with a larger number of parameters than the initial neural network model, the soft label is the operating state of the Internet of Things device in the historical operating period predicted by the teacher model, and the prediction result is the operating state of the Internet of Things device in the historical operating period predicted by the initial neural network model; Determine a first loss function according to the soft label and the prediction result, determine a second loss function according to the prediction result and the true label, and determine a total loss function according to the first loss function and the second loss function; The training process of the initial neural network model is determined according to the total loss function, wherein the training process includes: continuing iterative training and stopping iterative training.

6. The method according to claim 1, characterized in that During the training process, multiple neural network layers of the initial neural network model are deleted by the following method: Obtaining the absolute value of the weight of each layer in the initial neural network model, and arranging the multiple absolute values ​​in ascending order into a weight sequence; Determining a pruning threshold according to the weight sequence, wherein the pruning threshold is used to assist in determining a neural network layer to be deleted; Assigning an absolute value less than the pruning threshold as an invalid value, and assigning an absolute value greater than the pruning threshold as a valid value; The neural network layer corresponding to the invalid value is determined as the deleted neural network layer, and the iteration is stopped until the sparsity of the initial neural network model reaches the preset sparsity to obtain the abnormal probability analysis model, wherein the preset sparsity is used to indicate the ratio of the number of deleted neural network layers to the number of all neural network layers contained in the initial neural network model.

7. The method according to claim 6, characterized in that Determining a pruning threshold according to the weight sequence includes: During the first iterative training, determining a deletion ratio corresponding to each absolute value, and when the deletion ratio is equal to the preset sparsity, determining the absolute value corresponding to the deletion ratio as the pruning threshold, wherein the deletion ratio is used to indicate the ratio of the number of other absolute values ​​in the weight sequence whose values ​​are smaller than the absolute value to the total number of absolute values ​​contained in the weight sequence; During non-first iterative training, the weight sequence is updated according to the result of the previous iterative training to obtain a new weight sequence; the deletion ratio corresponding to each absolute value in the new weight sequence is determined, and when the deletion ratio is equal to the preset sparsity, the absolute value corresponding to the deletion ratio is determined as the current pruning threshold; the pruning threshold is determined according to the historical pruning threshold used in the previous iterative training process and the current pruning threshold.

8. The method according to claim 1, characterized in that Determining whether the device to be detected has an abnormality according to the plurality of probability matrices corresponding to the preset time window, and obtaining a determination result, including: For each of the probability matrices, determining a target element in the probability matrix that is greater than a preset probability value, and if the target element exists in any of the probability matrices, determining that the judgment result is that the device to be detected has an abnormality, and determining the abnormality type corresponding to the target element as the abnormality type of the device to be detected, the abnormality type including: network attack, performance bottleneck, and hardware failure; In a case where the target element does not exist in any of the plurality of probability matrices corresponding to the preset time window, it is determined that the judgment result is that there is no abnormality in the device to be detected.

9. The method according to claim 8, characterized in that Determine the alarm level of each abnormality type according to the multiple judgment results corresponding to the multiple preset time windows: Determining the number of alarms corresponding to each of the abnormality types, wherein the number of alarms is the number of times the abnormality type is recorded in the plurality of judgment results; Determine the alarm level of the abnormality type corresponding to the number of alarms greater than or equal to the first preset value as the first level; Determine the alarm level of the abnormality type corresponding to the number of alarms greater than or equal to a second preset value as a second level, wherein the second preset value is greater than the first preset value, and the second level is higher than the first level; The alarm level of the abnormality type corresponding to the number of alarms greater than or equal to a third preset value is determined as a third level, wherein the third preset value is greater than the second preset value, and the third level is higher than the second level.

10. A device for detecting cloud network equipment, characterized in that: include: an acquisition module, configured to acquire multiple operating data of a device to be detected during a detection period, and generate a multimodal feature fusion vector based on the multiple operating data, wherein the device to be detected is a cloud network device in a cloud network environment, and the number of the device to be detected is multiple; a processing module configured to process and analyze the multimodal feature fusion vector using an abnormality probability analysis model for each device to be detected, thereby obtaining a probability matrix output by the abnormality probability analysis model, wherein each element in the probability matrix indicates the probability of a class of abnormalities existing in the device to be detected, and the abnormality probability analysis model is an initial neural network model with multiple neural network layers deleted during training; a judgment module, configured to judge whether the device to be detected has an abnormality in the preset time window based on the plurality of probability matrices corresponding to the preset time window, and obtain a judgment result, wherein the preset time window includes a plurality of consecutive detection time periods, and the judgment result records the abnormality type of the device to be detected; A determination module is used to determine the alarm level of each abnormality type according to the multiple judgment results corresponding to the multiple preset time windows.

11. A non-volatile storage medium, characterized in that: The non-volatile storage medium stores a computer program, wherein the device where the non-volatile storage medium is located executes the method for detecting cloud network devices according to any one of claims 1 to 9 by running the computer program.

12. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to execute the method for detecting cloud network devices according to any one of claims 1 to 9 through the computer program.

13. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by the processor, the steps of the method for detecting cloud network devices described in any one of claims 1 to 9 are implemented.