A device monitoring method and system based on deep learning

Through the deep learning-based device monitoring method, the preset threshold of electronic devices is dynamically adjusted, and the problem of insufficient judgment of a single fixed threshold in the prior art is solved, thereby achieving more efficient fault warning and system stability.

CN119806967BActive Publication Date: 2025-05-30YUNBIAN CLOUD TECHNOLOGY (SHANGHAI) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510308616.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-05-30
Estimated Expiration
2045-03-17

AI Technical Summary

Technical Problem

The prior art relies on a single fixed threshold for fault judgment in electronic equipment monitoring, resulting in the inability to detect potential faults in time and affect system stability.

Method used

Using deep learning-based device monitoring method, the Agent module uses eBPF technology to obtain the operating data of electronic devices in real time, evaluate the equipment importance level, and dynamically adjust the preset threshold according to the associated anomaly index to output anomaly alarm signal.

Benefits of technology

It improves the accuracy and timeliness of fault warning, reduces the probability of equipment damage, improves the stability and reliability of the system, and reduces business interruptions and economic losses caused by equipment failure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119806967B_ABST
    Figure CN119806967B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of computer technology, and provides a device monitoring method and system based on deep learning. The method includes: the Agent module at the monitoring end obtains the first real-time operation data of each electronic device to be monitored, and determines the electronic devices that exceed the corresponding risk thresholds as target electronic devices; evaluates the importance level of the target electronic devices. If the importance level is higher than the level threshold, obtains and evaluates the correlation anomaly index based on the second real-time operation data of each associated electronic device in the same geographical area as the target electronic device and the corresponding risk thresholds, and lowers the risk threshold of the target electronic device to a second preset threshold; if the first real-time operation data of the target electronic device exceeds the second preset threshold, outputs an abnormal alarm signal for the target electronic device. The present invention takes into account the correlation between electronic devices and surrounding environmental factors, and can improve the accuracy and timeliness of fault warning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to a device monitoring method and system based on deep learning. Background Art

[0002] In the current field of electronic device monitoring, with the continuous increase in the number of electronic devices and the increasing complexity of application scenarios, effective monitoring of the operating status of electronic devices has become particularly important. Currently, the monitoring end establishes communication connections with multiple electronic devices to receive real-time operating data of each electronic device, and then uses anomaly analysis algorithms to analyze and judge these data. When the real-time operating data exceeds a preset threshold, it is determined that the device is abnormal, that is, a fault may have occurred.

[0003] However, the existing technology has obvious defects. The currently designed preset thresholds are usually corresponding to each electronic device one by one, and this method only judges whether a fault exists based on the preset threshold of a single electronic device itself; moreover, in order to avoid misjudgment, the preset threshold is usually designed to be slightly higher, and too high a preset threshold may cause the electronic device not to be alarmed and disposed of in time for faults, which is not conducive to the health of electronic devices, especially those with a higher degree of importance, resulting in poor system stability.

[0004] Therefore, how to dynamically adjust the preset threshold of electronic devices to achieve more effective monitoring and fault early warning processing of electronic devices is a technical problem that needs to be solved urgently at present. Summary of the Invention

[0005] In view of the above technical problems, the present invention provides a device monitoring method, system, electronic device, computer storage medium, and computer program product based on deep learning.

[0006] The present invention discloses a device monitoring method based on deep learning, and the method includes the following steps:

[0007] The Agent module of the monitoring end uses eBPF technology to obtain the first real-time operating data of each monitored electronic device in real time, and determines the electronic device whose first real-time operating data exceeds the corresponding risk threshold as the target electronic device; wherein, the risk threshold is lower than the standard preset threshold.

[0008] Evaluate the importance level of the target electronic device. If the importance level is higher than the level threshold, obtain the second real-time operating data of each associated electronic device in the same geographical area as the target electronic device.

[0009] An associated anomaly index is evaluated based on each of the second real-time operation data and the corresponding risk threshold, and the standard preset threshold of the target electronic device is adjusted to a temporary standard preset threshold according to the associated anomaly index;

[0010] If the first real-time operation data of the target electronic device exceeds the temporary standard preset threshold, an anomaly alarm signal for the target electronic device is output.

[0011] Optionally, evaluating the importance level of the target electronic device includes:

[0012] Obtain the performance parameters of the target electronic device itself, and evaluate a first importance score and a second importance score based on the performance parameters; wherein, the operation record data includes data that can be used to evaluate the operation reliability of the target electronic device;

[0013] Obtain the function link information to which the target electronic device belongs in the entire system where it is deployed, and evaluate a third importance score based on the function link information;

[0014] Perform weighted fusion on the first importance score, the second importance score, and the third importance score to obtain a fourth importance score, and match the importance level according to the fourth importance score.

[0015] Optionally, obtaining the second real-time operation data of each associated electronic device in the same geographical area as the target electronic device includes:

[0016] Match a first screening distance from the pre-set first associated data according to the importance level of the target electronic device;

[0017] Obtain all historical fault records of the target electronic device, classify each of the historical fault records into an internal cause fault type and an external cause fault type in turn, count the type quantity of the external cause fault type, and match an increased distance weight coefficient from the pre-set second associated data according to the type quantity;

[0018] Use the increased distance weight coefficient to adjust the first screening distance to a larger second screening distance.

[0019] Optionally, evaluating the associated anomaly index based on each of the second real-time operation data and the corresponding risk threshold includes:

[0020] Calculate a first difference between each of the second real-time operation data of each associated electronic device and the corresponding risk threshold, and calculate an equivalent value of the normalized values of all the first differences;

[0021] If the equivalent value is higher than the preset value, further calculate the second difference between the first real-time operation data of the target electronic device and the corresponding risk threshold, and obtain the first associated anomaly index according to the normalized value of the second difference and the equivalent value;

[0022] If the equivalent value is not higher than the preset value, further obtain the respective historical operation data of each associated electronic device in the near future, fit the operation trend curve according to the historical operation data and the second real-time operation data of the same type, and identify the abnormal trend curve segment therein;

[0023] Use a convolutional network to extract features from each of the vectorized abnormal trend curve segments to obtain the operation sub-features of any associated electronic device, splice and integrate each of the operation sub-features into operation features, and use an associated anomaly prediction model to perform predictive analysis on the operation features, the first real-time operation data of the target electronic device, and the corresponding risk threshold to obtain the second associated anomaly index.

[0024] Optionally, the Agent module regularly collects and updates the historical fault records of the target electronic device, constructs a training data set using the historical fault records, and locally trains the associated anomaly prediction model using the training data in the training data set;

[0025] In addition, the Agent module also regularly publishes the associated anomaly prediction model to the consortium blockchain network for distributed training, integrates the important parameters of the model obtained from the distributed training into the locally trained associated anomaly prediction model, and then performs secondary training.

[0026] The present invention also discloses a device monitoring system based on deep learning. The system includes a processing device and a storage device. The computer code stored in the storage device is called and executed by the processing device to implement the following steps:

[0027] The Agent module at the monitoring end uses the eBPF technology to obtain the first real-time operation data of each monitored electronic device in real time, and determines the electronic device whose first real-time operation data exceeds the corresponding risk threshold as the target electronic device; wherein, the risk threshold is lower than the standard preset threshold;

[0028] Evaluate the importance level of the target electronic device. If the importance level is higher than the level threshold, obtain the second real-time operation data of each associated electronic device in the same geographical area as the target electronic device;

[0029] Evaluate and obtain the associated anomaly index according to each of the second real-time operation data and the corresponding risk threshold, and adjust the standard preset threshold of the target electronic device to a temporary standard preset threshold according to the associated anomaly index;

[0030] If the first real-time operation data of the target electronic device exceeds the temporarily preset threshold, an abnormal alarm signal for the target electronic device is output.

[0031] The present invention also discloses an electronic device, including: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, where the processor executes the computer program to implement the method as described in any one of the foregoing.

[0032] The present invention also discloses a computer storage medium, where the computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the method as described in any one of the foregoing.

[0033] The present invention also discloses a computer program product, where the computer program product contains computer code, and when the computer code is executed by a processor of an electronic device, the method as described in any one of the foregoing is implemented.

[0034] The beneficial effects of the present invention are at least as follows:

[0035] Compared with the traditional single fixed threshold judgment method, the above method of the present invention can more comprehensively consider the correlation between electronic devices and surrounding environmental factors, greatly improving the accuracy and timeliness of fault warning. For important electronic devices, potential fault risks can be discovered in advance, the probability of equipment damage can be effectively reduced, the stability and reliability of the entire system are improved, business interruptions and economic losses caused by equipment failures are reduced, and a strong guarantee is provided for the stable operation of electronic devices. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0037] Figure 1 is a flowchart of a device monitoring method based on deep learning disclosed in an embodiment of the present invention;

[0038] Figure 2 is a structural diagram of a device monitoring system based on deep learning disclosed in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0039] The following specific embodiments illustrate the implementation manners of the present application. Those skilled in the art can easily understand the other advantages and effects of the present application from the content disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts belong to the scope of protection of the present application.

[0040] In addition, the technical features involved in different implementation manners of the present application described below can be combined with each other as long as they do not conflict with each other.

[0041] In response to the above technical problems, as Figure 1 shown, an embodiment of the present invention discloses a device monitoring method based on deep learning, and the method includes the following steps:

[0042] S101, the Agent module at the monitoring end uses the eBPF technology to obtain the first real-time operation data of each monitored electronic device in real time, and determines the electronic device whose first real-time operation data exceeds the corresponding risk threshold as the target electronic device; wherein, the risk threshold is lower than the standard preset threshold.

[0043] The Agent module at the monitoring end uses the eBPF (extended Berkeley Packet Filter) technology to obtain the first real-time operation data of each monitored electronic device in real time. The eBPF technology is an efficient kernel-state programmable technology, which can deeply penetrate into the operating system kernel without affecting the system performance and accurately capture various system data. For example, in a large cloud computing data center with thousands of servers, the Agent module can use the eBPF technology to obtain the first real-time operation data such as the CPU usage rate, memory occupancy rate, and disk I / O rate of each server in real time. In addition, the Agent module also supports the OTLP (Open Telemetry Protocol) and Prometheus protocols, so as to be able to collect millions of time-series data with extremely low system overhead.

[0044] If any first real-time operation data of a certain electronic device does not exceed the standard preset threshold but exceeds the corresponding risk threshold, it is determined that the electronic device has an abnormal tendency, and at this time it is determined as the target electronic device. For example, if the CPU usage rate of a certain server exceeds the threshold of 80% (but is lower than the standard preset threshold of 90%), then this server will be determined as the target electronic device.

[0045] S102. Evaluate the importance level of the target electronic device. If the importance level is higher than the level threshold, obtain the second real-time operation data of each associated electronic device in the same geographical area as the target electronic device.

[0046] The evaluation of the importance level can be determined based on factors such as the function of the electronic device in the entire system and the criticality of the business it undertakes. For example, in a financial trading system, the server responsible for core transaction processing has a much higher importance level than the server responsible for log recording. If the importance level is higher than the level threshold, it indicates that the target electronic device is very critical, and once a failure occurs, it will have a serious impact on the entire system. At this time, in order to more comprehensively and accurately judge the status of the target electronic device, it is necessary to obtain the second real-time operation data of each associated electronic device in the same geographical area as the target electronic device. For example, in the same computer room of a data center, if a key database server is determined as the target electronic device, then it is necessary to obtain the real-time operation data of other servers, network devices and other associated electronic devices in the same computer room, such as the traffic data of network devices and the load data of other servers.

[0047] S103. Evaluate and obtain an associated anomaly index based on each of the second real-time operation data and the corresponding risk threshold, and adjust the standard preset threshold of the target electronic device to a temporary standard preset threshold according to the associated anomaly index.

[0048] In this step, by comparing the real-time operation data of the associated electronic devices with their respective risk thresholds (which are also lower than the standard preset threshold), an associated anomaly index is comprehensively analyzed to measure the synchronization anomaly degree between these associated electronic devices and the target electronic device. For example, if the CPU usage rates of multiple servers in the same computer room are all close to or exceed their own risk thresholds (but are all lower than the standard preset threshold), and the traffic of network devices also far exceeds the normal level; or, there is strong electromagnetic interference in the computer room, and the data transmission rates of multiple servers in the computer room significantly decrease. In the above cases, the corresponding associated anomaly index will be relatively high.

[0049] If the above associated anomaly index is very high, it indicates that there are relatively large risks in the environment where the target electronic device is located. In order to timely detect possible failures of the target electronic device, it is necessary to adjust the standard preset threshold of the target electronic device. For example, originally the standard preset threshold of the CPU usage rate of the target electronic device is 90%. When the associated anomaly index shows that there are also obvious anomalies in the surrounding devices, the threshold can be set to be lowered to 80%; or, originally the standard preset threshold (lower limit value) of the data transmission rate of the target electronic device is 50 GT / s. When the associated anomaly index shows that there are also obvious anomalies in the surrounding devices, the threshold can be set to be raised to 60 GT / s.

[0050] S104, if the first real-time operation data of the target electronic device exceeds the temporarily set standard preset threshold, an abnormal alarm signal for the target electronic device is output.

[0051] After the previous steps, the threshold of the target electronic device is adjusted more reasonably. At this time, once its first real-time operation data (for example, the one with the largest amplitude exceeding the corresponding risk threshold) exceeds (that is, is higher or lower than) the adjusted temporarily set standard preset threshold, it can be determined that the target electronic device is abnormal, and an alarm signal is output in a timely manner. For example, when the CPU usage rate of the target electronic device exceeds the lowered threshold of 80%, or the data transmission rate is lower than the raised threshold of 60 GT / s, the monitoring system will immediately issue an alarm to notify the operation and maintenance personnel to handle it, avoiding the further deterioration of the fault. In addition, after the alarm is completed, the temporarily set standard preset threshold is restored to the standard preset threshold.

[0052] Compared with the traditional single fixed threshold judgment method, the above method of the present invention can more comprehensively consider the relevance between electronic devices and surrounding environmental factors, greatly improving the accuracy and timeliness of fault warning. For important electronic devices, potential fault risks can be discovered in advance, the probability of equipment damage can be effectively reduced, the stability and reliability of the entire system are improved, business interruptions and economic losses caused by equipment failures are reduced, providing a strong guarantee for the stable operation of electronic devices.

[0053] Optionally, evaluating the importance level of the target electronic device includes:

[0054] Obtain the performance parameters of the target electronic device itself, and evaluate the first importance score and the second importance score according to the performance parameters; wherein, the operation record data includes data that can be used to evaluate the operation reliability of the target electronic device.

[0055] Obtain the functional link information to which the target electronic device belongs in the entire system where it is deployed, and evaluate the third importance score according to the functional link information.

[0056] Perform weighted fusion on the first importance score, the second importance score, and the third importance score to obtain the fourth importance score, and match the importance level according to the fourth importance score.

[0057] In this embodiment, the performance parameters of the target electronic device itself are one of the key indicators for measuring its importance. For example, in a data center, parameters such as the number of CPU cores, memory capacity, and hard disk read / write speed of a server determine its task processing ability. A high-performance server can quickly respond to a large number of business requests. If its performance deteriorates or a failure occurs, the impact on the entire system will be significant. Through specific evaluation algorithms (such as the Analytic Hierarchy Process AHP, Multilayer Perceptron MLP, etc.), these performance parameters are quantified into the first importance score, and the better the parameters, the higher the score.

[0058] To control costs, for important key components, designers generally choose equipment produced based on higher product design and production standards, while for those unimportant non-key components, they will choose equipment produced based on lower product design and production standards. Therefore, by statistically analyzing the number of failures, mean time between failures, etc. in the operation record data of the target electronic device, its failure rate is evaluated, and based on the level of the failure rate, it is determined whether it belongs to a key device or a non-key device. If it is statistically found that a certain device has a high failure rate, it indicates that it is probably a non-key device and has a lower importance in the entire production system, and then the second importance score evaluated for it is set to a lower value; conversely, if the failure rate of a certain electronic device is very low, it indicates that it is very important and is probably a key device, and then the second importance score evaluated for it is set to a higher value.

[0059] The information about the functional link to which the target electronic device belongs in the entire system where it is deployed is also crucial for judging its importance level. For example, in an e-commerce platform system, the server responsible for order processing is in the core functional link, which is directly related to the completion of transactions and has a higher importance; while the server responsible for displaying user comments, although it also has a certain role, has a lower importance compared to the order processing link. By analyzing the criticality of the functional link in the entire system, the third importance score is evaluated, and the more critical the function, the higher the score.

[0060] The first importance score, second importance score, and third importance score obtained above are weighted and fused. Different score weights can be set according to the actual situation and experience. For example, if the system pays more attention to the performance of the device, the weight of the first importance score can be set higher; if it values the operational reliability of the device more, the weight of the second importance score can be increased. Through weighted calculation, the fourth importance score is obtained, which comprehensively reflects various important factors of the device.

[0061] Finally, according to the pre-set matching rules between the score and the importance level, for example, the importance scores are divided into different intervals, and each interval corresponds to a different importance level (high, medium, low, etc.), so as to match and obtain the importance level of the target electronic device.

[0062] Optionally, obtaining the second real-time operation data of each associated electronic device in the same geographical area as the target electronic device includes:

[0063] Matching a first screening distance from the pre-set first associated data according to the importance level of the target electronic device;

[0064] Obtaining all historical fault records of the target electronic device, classifying each of the historical fault records into an internal cause fault type and an external cause fault type in sequence, counting the number of types of the external cause fault type, and matching an increased distance weight coefficient from the pre-set second associated data according to the number of types;

[0065] Adjusting the first screening distance to a larger second screening distance using the increased distance weight coefficient.

[0066] In this embodiment, the abnormal situation of the target electronic device may be caused by external factors (such as electromagnetic interference), and external factors usually have an impact on other electronic devices in the area at the same time. At this time, it can be assisted to analyze whether there is an external cause abnormality in the target electronic device by referring to whether other electronic devices in the same geographical area (such as the same server room) also generally show abnormalities. For the scope of the geographical area, it can be determined in the following way:

[0067] Pre-set first associated data for characterizing the corresponding relationship between the device importance level and the screening distance, and then obtaining the matched first screening distance from this first associated data according to the evaluated importance level of the target electronic device. For example, for a target electronic device with a high importance level, a larger first screening distance will be matched in the first associated data. For example, in a large data center, the core server has a high importance level, and its first screening distance may be set to the associated electronic devices within a radius of 20 meters centered on the server; while for a non-core server with a lower importance level, the screening distance can be a radius of 10 meters.

[0068] At the same time, different target electronic devices have different abilities to resist the influence of external factors. By analyzing the number of faults of the external cause type in all historical fault records of the target electronic device, it can be used to indirectly analyze this ability. Specifically:

[0069] Obtain all historical fault records of the target electronic device and classify them into internal cause fault types and external cause fault types. Internal cause faults are usually caused by internal factors such as equipment overload and software errors; external cause faults are caused by external factors such as electromagnetic interference, power supply abnormalities, and physical vibrations.

[0070] Count the number of external cause failure types, which reflects the degree to which the target electronic device is vulnerable to the external environment. Similar to the first associated data, the preset second associated data is used to characterize the corresponding relationship between the number of external cause failure types and the distance-increasing weight coefficient. If the number of external cause failure types is large, it indicates that the target electronic device has a smaller ability to resist external influences, that is, multiple external causes are likely to cause the target electronic device to fail, and the distance-increasing weight coefficient matched from the second associated data is large; otherwise, it is small. For example, if the target electronic device has failed due to 10 external causes such as electromagnetic interference and voltage fluctuations in history, the matched distance-increasing weight coefficient is 1.5; if the number of external cause failure types is 5, the distance-increasing weight coefficient is set to 1.1.

[0071] Next, use the obtained distance-increasing weight coefficient to adjust the first screening distance to a larger second screening distance. In this way, it is possible to decide to what extent to increase the first screening distance based on the degree of difficulty of the target electronic device being affected by external factors. When the target electronic device is more vulnerable to multiple external environments, it may not be comprehensive enough to obtain associated electronic device data only based on the first screening distance (for example, some electronic devices within the first screening distance are not affected by electromagnetic interference, and at this time these electronic devices cannot be used to assist in analyzing whether the target electronic device is affected by electromagnetic interference). It is necessary to expand the scope to obtain the real-time operation data of more associated electronic devices that may be affected by the same external factors, so as to improve the accuracy and timeliness of the abnormal analysis of the target electronic device.

[0072] It should be noted that the historical failure record refers to the failure cause record data obtained by manual analysis of the faulty electronic device by the management personnel, or the failure cause record data automatically analyzed and recorded by the management system based on the monitored failure data.

[0073] Optionally, the evaluating the associated abnormal index according to each of the second real-time operation data and the corresponding risk threshold includes:

[0074] Calculate the first difference between each of the second real-time operation data of each associated electronic device and the corresponding risk threshold, and calculate the equivalent value of the normalized values of all the first differences;

[0075] If the equivalent value is higher than the preset value, further calculate the second difference between the first real-time operation data of the target electronic device and the corresponding risk threshold, and obtain the first associated abnormal index according to the normalized value of the second difference and the equivalent value;

[0076] If the equivalent value is not higher than the preset value, further obtain the recent historical operation data of each associated electronic device, fit the operation trend curve according to the historical operation data and the second real-time operation data of the same type, and identify the abnormal trend curve segment therein;

[0077] Use a convolutional network to extract features from each of the vectorized abnormal trend curve segments to obtain the operating sub-features of any associated electronic device, splice and integrate each of the operating sub-features into operating features, and use an associated anomaly prediction model to perform predictive analysis on the operating features, the first real-time operating data of the target electronic device, and the corresponding risk thresholds to obtain a second associated anomaly index.

[0078] In this embodiment, multiple pieces of second real-time operating data of each associated electronic device are obtained (each piece of second real-time operating data corresponds to different components or modules of the associated electronic device, or different operating parameters of the same component or module), and the first difference (risk threshold - second real-time operating data) between each piece of second real-time operating data and the corresponding risk threshold is calculated. For example, if the second real-time operating data of the CPU usage rate of an associated electronic device is 65%, and its corresponding risk threshold is 70%, then the first difference is 70% - 65% = 5%. After calculating the multiple first differences of each associated electronic device, the equivalent value of the normalized values of these first differences (eliminating the influence of the dimensions of different operating parameters) is calculated. The equivalent value can be the median value or the maximum value among the normalized values of the first differences, etc., and is not specifically limited.

[0079] When the calculated equivalent value is higher than the preset value, it indicates that an associated electronic device has a trend of significantly deviating from the normal range, which indirectly reflects that the associated electronic device and the target electronic device are likely to have abnormal conditions together. At this time, the above equivalent value of the associated electronic device can be directly used to calculate the associated anomaly index of the target electronic device: First, calculate the second difference (first real-time operating data - risk threshold) between the first real-time operating data of the target electronic device (for example, the one with the largest amplitude exceeding the corresponding risk threshold) and the corresponding risk threshold; then, calculate (for example, calculate the difference between the two) the first associated anomaly index based on the normalized value of this second difference and the previously calculated equivalent value. The first associated anomaly index and the difference between the two are in a negative correlation.

[0080] If the equivalent value is not higher than the preset value, it indicates that the deviation degree of the current operating data of the associated electronic device from the preset threshold is within an acceptable range, without obvious anomalies but still potentially abnormal. At this time, the historical operating data of each associated electronic device in the recent period (for example, within 20s) is obtained, and these historical operating data are fitted with the second real-time operating data of the same type (for example, the same data transmission rate) to generate an operating trend curve. Abnormal trend curve segments are identified through data analysis methods, such as a sudden sharp drop in the curve of the data transmission rate.

[0081] Vectorize the identified abnormal trend curve segments so that they can be processed by a convolutional network. The convolutional network has powerful feature extraction capabilities. By performing operations such as convolutional operations and pooling on the vectorized abnormal trend curve segments, the operating sub-features of any associated electronic device are extracted. For example, features such as the frequency and amplitude of curve changes are extracted. Then, the operating sub-features are spliced and integrated into operating features, which comprehensively reflect the abnormal conditions of all associated electronic devices. Finally, use the pre-trained associated abnormal prediction model to perform predictive analysis on the operating features, together with the first real-time operating data of the target electronic device and the corresponding risk threshold, to obtain the second associated abnormal index. Among them, multiple abnormal trend curve segments corresponding to different operating parameters can be obtained, and the maximum value among the predicted and analyzed associated abnormal indexes corresponding to these abnormal trend curve segments is used as the second associated abnormal index.

[0082] Among them, the associated abnormal prediction model can be constructed based on machine learning algorithms (such as support vector machines, random forests, etc.) or deep learning algorithms (such as recurrent neural networks, etc.). Through learning a large amount of historical data, it can predict the abnormal degree of the current associated electronic device according to the input operating features, the first real-time operating data of the target electronic device, and the corresponding risk threshold. However, the associated abnormal prediction model can also be obtained by locally fine-tuning and training a general large model (such as DeepSeek, GPT-4), and the specific method is not limited.

[0083] Optionally, the Agent module regularly collects and updates the historical fault records of the target electronic device, constructs a training data set using the historical fault records, and locally trains the associated abnormal prediction model using the training data in the training data set;

[0084] In addition, the Agent module also regularly publishes the associated abnormal prediction model to the alliance chain network for distributed training, integrates the important parameters of the model obtained from the distributed training into the locally trained associated abnormal prediction model, and then performs secondary training.

[0085] In this embodiment, during operation, the Agent module also undertakes the key task of regularly collecting and updating the historical fault records of the target electronic device. These historical fault records contain various fault information that occurred during the past operation of the device, such as the time of fault occurrence, fault type, and operating parameters of the device at the time of fault. By sorting and screening these rich and detailed data, the Agent module can construct a training data set for model training. For example, data with the same fault type are grouped into one category, and at the same time, the key operating parameters related to the fault are extracted as features, thus forming training data with clear features and labels.

[0086] Using the constructed training dataset, the Agent module first locally trains the associated anomaly prediction model using the training data. Through repeated iterative training, the model is gradually optimized and can more accurately predict whether there are associated anomalies based on the input device operation data.

[0087] The Agent module also has the ability to publish the associated anomaly prediction model to the consortium blockchain network for distributed training. The consortium blockchain network consists of multiple nodes, and each node can use its own computing resources and data to train the model. This distributed training method can make full use of the resources and data of all parties, increase the diversity and scale of training data, and thus improve the generalization ability of the model. After the distributed training is completed, the Agent module will integrate the important parameters of the models obtained by each node's training (different according to the type of the model, such as the learning rate, the number of convolutional kernels of the CNN model, etc.) into the locally trained associated anomaly prediction model. For example, the prediction parameters for a specific fault type in the models obtained by different nodes are weighted and averaged, and then updated to the local model.

[0088] After completing the parameter integration, the Agent module will conduct secondary training on the model to fine-tune the newly integrated parameters so that they can work better with the existing local parameters, and thus can more accurately predict associated anomalies when facing complex and changing device operation conditions.

[0089] As Figure 2 shown, an embodiment of the present invention also discloses a device monitoring system based on deep learning. The system includes a processing device and a storage device. The computer code stored in the storage device is called and executed by the processing device to implement the following steps:

[0090] The Agent module at the monitoring end uses the eBPF technology to obtain the first real-time operation data of each monitored electronic device in real time, and determines the target electronic device as the electronic device whose first real-time operation data exceeds the corresponding risk threshold; wherein, the risk threshold is lower than the standard preset threshold;

[0091] Evaluate the importance level of the target electronic device. If the importance level is higher than the level threshold, obtain the second real-time operation data of each associated electronic device in the same geographical area as the target electronic device;

[0092] Evaluate the associated anomaly index according to each second real-time operation data and the corresponding risk threshold, and adjust the standard preset threshold of the target electronic device to a temporary standard preset threshold according to the associated anomaly index;

[0093] If the first real-time operation data of the target electronic device exceeds the temporary standard preset threshold, an abnormal alarm signal for the target electronic device is output.

[0094] An embodiment of the present invention also discloses an electronic device, including: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, where the processor executes the computer program to implement the method as described in the foregoing embodiment.

[0095] An embodiment of the present invention also discloses a computer storage medium storing a computer program, where the computer program is executed by a processor to implement the method as described in the foregoing embodiment.

[0096] An embodiment of the present invention also discloses a computer program product containing computer code, where when the computer code is executed by a processor of an electronic device, the method as described in the foregoing embodiment is implemented.

[0097] The above-mentioned computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or equipment, or any suitable combination of the above. Alternatively, the computer-readable storage medium may be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above.

[0098] It should be understood that various forms of the processes shown above may be used, with steps reordered, added, or deleted. For example, the steps described in the present invention may be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.

[0099] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A device monitoring method based on deep learning, characterized in that: The method comprises the following steps: The Agent module of the monitoring end uses the eBPF technology to obtain the first real-time operation data of each monitored electronic device in real time, and determines the electronic device whose first real-time operation data exceeds the corresponding risk threshold as the target electronic device; wherein the risk threshold is lower than the standard preset threshold; Assessing the importance level of the target electronic device, and if the importance level is higher than a level threshold, acquiring second real-time operation data of each associated electronic device in the same geographical area as the target electronic device; Determine an associated anomaly index based on each of the second real-time operation data and the corresponding risk threshold, and adjust the standard preset threshold of the target electronic device to a temporary standard preset threshold based on the associated anomaly index; the associated anomaly index is used to measure the degree of synchronization anomaly between each associated electronic device and the target electronic device; If the first real-time operation data of the target electronic device exceeds the temporary standard preset threshold, outputting an abnormal alarm signal for the target electronic device; The step of evaluating the associated abnormality index based on each of the second real-time operation data and the corresponding risk threshold comprises: Calculating a first difference between each of the second real-time operating data of each associated electronic device and a corresponding risk threshold, and calculating an equivalent value of a normalized value of all the first differences; If the equivalent value is higher than a preset value, further calculating a second difference between the first real-time operation data of the target electronic device and the corresponding risk threshold, and obtaining a first correlation abnormality index according to a normalized value of the second difference and the equivalent value; If the equivalent value is not higher than the preset value, further acquiring each recent historical operation data of each associated electronic device, fitting the historical operation data with the second real-time operation data of the same type to obtain an operation trend curve, and identifying an abnormal trend curve segment therein; A convolutional network is used to extract features from the vectorized abnormal trend curve segments to obtain operation sub-features of any associated electronic device, and the operation sub-features are spliced ​​and integrated into operation features. An associated abnormality prediction model is used to predict and analyze the operation features and the first real-time operation data of the target electronic device and the corresponding risk threshold to obtain a second associated abnormality index.

2. The device monitoring method based on deep learning according to claim 1, characterized in that: The evaluation target electronic device importance level includes: Acquire performance parameters of the target electronic device itself, and quantify the performance parameters to obtain a first important score; acquire operation record data of the target electronic device, and evaluate the operation record data to obtain a second important score; wherein the operation record data includes data that can be used to evaluate the operation reliability of the target electronic device; Acquire functional link information of the target electronic device in the entire deployed system, and evaluate the criticality of the functional link information in the entire system to obtain a third importance score; The first importance score, the second importance score, and the third importance score are weightedly integrated to obtain a fourth importance score, and the importance level is obtained based on matching of the fourth importance score.

3. The device monitoring method based on deep learning according to claim 2, characterized in that: The step of obtaining the second real-time operation data of each associated electronic device in the same geographical area as the target electronic device includes: According to the importance level of the target electronic device, a first screening distance is obtained by matching the preset first association data; Acquire all historical fault records of the target electronic device, classify each of the historical fault records into an internal fault type and an external fault type, obtain the number of external fault types by counting, and obtain a range-increasing weight coefficient by matching from a preset second associated data according to the number of types; The first screening distance is adjusted to a larger second screening distance using the range-increasing weighting factor.

4. The device monitoring method based on deep learning according to claim 1, characterized in that: The Agent module regularly collects and updates historical fault records of the target electronic device, constructs a training data set using the historical fault records, and locally trains the associated anomaly prediction model using the training data in the training data set; In addition, the Agent module also regularly publishes the associated anomaly prediction model to the alliance chain network for distributed training, and integrates the important parameters of the model obtained from the distributed training into the associated anomaly prediction model after local training, and then performs secondary training.

5. A deep learning-based equipment monitoring system, the system comprising a processing device and a storage device, characterized in that: The computer code stored in the storage device is called and executed by the processing device to implement the following steps: The Agent module of the monitoring end uses the eBPF technology to obtain the first real-time operation data of each monitored electronic device in real time, and determines the electronic device whose first real-time operation data exceeds the corresponding risk threshold as the target electronic device; wherein the risk threshold is lower than the standard preset threshold; Assessing the importance level of the target electronic device, and if the importance level is higher than a level threshold, acquiring second real-time operation data of each associated electronic device in the same geographical area as the target electronic device; Determine an associated anomaly index based on each of the second real-time operation data and the corresponding risk threshold, and adjust the standard preset threshold of the target electronic device to a temporary standard preset threshold based on the associated anomaly index; the associated anomaly index is used to measure the degree of synchronization anomaly between each associated electronic device and the target electronic device; If the first real-time operation data of the target electronic device exceeds the temporary standard preset threshold, outputting an abnormal alarm signal for the target electronic device; The step of evaluating the associated abnormality index based on each of the second real-time operation data and the corresponding risk threshold comprises: Calculating a first difference between each of the second real-time operating data of each associated electronic device and a corresponding risk threshold, and calculating an equivalent value of a normalized value of all the first differences; If the equivalent value is higher than a preset value, further calculating a second difference between the first real-time operation data of the target electronic device and the corresponding risk threshold, and obtaining a first correlation abnormality index according to a normalized value of the second difference and the equivalent value; If the equivalent value is not higher than the preset value, further acquiring each recent historical operation data of each associated electronic device, fitting the historical operation data with the second real-time operation data of the same type to obtain an operation trend curve, and identifying an abnormal trend curve segment therein; A convolutional network is used to extract features from the vectorized abnormal trend curve segments to obtain operation sub-features of any associated electronic device, and the operation sub-features are spliced ​​and integrated into operation features. An associated abnormality prediction model is used to predict and analyze the operation features and the first real-time operation data of the target electronic device and the corresponding risk threshold to obtain a second associated abnormality index.

6. The deep learning-based equipment monitoring system according to claim 5, characterized in that: The evaluation target electronic device importance level includes: Acquire performance parameters of the target electronic device itself, and quantify the performance parameters to obtain a first important score; acquire operation record data of the target electronic device, and evaluate the operation record data to obtain a second important score; wherein the operation record data includes data that can be used to evaluate the operation reliability of the target electronic device; Acquire functional link information of the target electronic device in the entire deployed system, and evaluate and obtain a third importance score based on the functional link information; The first importance score, the second importance score, and the third importance score are weightedly integrated to obtain a fourth importance score, and the importance level is obtained based on matching of the fourth importance score.

7. An electronic device comprising: At least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, wherein the processor executes the computer program to implement the method according to any one of claims 1 to 4.

8. A computer storage medium storing a computer program, characterized in that: The computer program is executed by a processor to implement the method according to any one of claims 1 to 4.

9. A computer program product, characterized in that: The computer program product includes computer codes, and when the computer codes are executed by a processor of an electronic device, the method according to any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • An alarm setting method and system based on system data monitoring

    CN109766247A

  • A cloud anomaly detection device using explainable ai based on deep learning and a anomaly detection method

    KR102698085B1