Data monitoring method and device, equipment and medium
By introducing distributed data acquisition servers and artificial intelligence algorithms into the cloud system, the problem of low data monitoring efficiency of different devices was solved, enabling real-time monitoring and anomaly detection of different professional software and computer terminals, thus improving the stability and security of the system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PETROCHINA CO LTD
- Filing Date
- 2024-11-04
- Publication Date
- 2026-05-08
AI Technical Summary
Existing technologies cannot effectively monitor data from different types of specialized software and different computer terminals, resulting in low efficiency in data sharing and monitoring.
By introducing a distributed data acquisition server into the cloud system, using preset data acquisition methods and environmental data acquisition methods, and combining artificial intelligence and deep learning algorithms for data integration and visualization, real-time monitoring and anomaly detection of the operating status and environmental data of different devices can be achieved.
It enables effective monitoring of data from various professional software and computer terminals, improving system stability and security, and allowing for timely detection and rapid response to equipment anomalies.
Smart Images

Figure CN121996716A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of big data monitoring technology, and in particular to a data monitoring method, apparatus, equipment, and medium. Background Technology
[0002] The exploration and development of unconventional natural gas refers to the exploration and development of natural gas reservoirs whose underground occurrence and accumulation patterns differ significantly from conventional natural gas reservoirs. During exploration and development, a large amount of data needs to be collected and professionally processed and analyzed. Related technologies mainly rely on cloud platforms for data resource sharing. However, because this data originates from different types of specialized software and different computer terminals, effective monitoring of this data is impossible. Therefore, how to effectively monitor data from different specialized software and different computer terminals is a pressing problem that needs to be solved. Summary of the Invention
[0003] This application provides a data monitoring method, apparatus, device, and medium, which solves the technical problem in the prior art that it is impossible to effectively monitor data from different types of professional software and different computer terminals, and achieves the technical effect of effectively monitoring data from different professional software and different computer terminals.
[0004] Firstly, this application provides a data monitoring method applied to a cloud system, wherein the cloud system communicates with different types of monitored devices, and the method includes:
[0005] According to the preset data collection methods corresponding to each monitored device, the operating status data of each monitored device is collected; each monitored device includes network devices, storage devices, server devices, middleware databases, big data clusters, and container clusters;
[0006] According to the environmental data collection method corresponding to each monitored device, the basic environmental data of each monitored device is collected;
[0007] The basic environmental data and the operating status data of each monitored device are integrated to obtain monitoring fusion data, which is then visualized.
[0008] Secondly, this application provides a data monitoring device applied to a cloud system, the cloud system communicating with different types of monitored devices, the device comprising:
[0009] The data acquisition module is used to collect the operating status data of each monitored device according to the preset data acquisition method corresponding to each monitored device; each monitored device includes network devices, storage devices, server devices, middleware databases, big data clusters, and container clusters.
[0010] The data acquisition module is used to collect basic environmental data of each monitored device according to the environmental data acquisition method corresponding to each monitored device.
[0011] The integrated display module is used to integrate basic environmental data and the operating status data of each monitored device to obtain integrated monitoring data, and to visualize the integrated monitoring data.
[0012] Thirdly, this application provides an electronic device, comprising:
[0013] processor;
[0014] Memory used to store processor-executable instructions;
[0015] The processor is configured to execute a data monitoring method as provided in the first aspect.
[0016] Fourthly, this application provides a non-transitory computer-readable storage medium that, when the instructions in the storage medium are executed by the processor of an electronic device, enables the electronic device to perform a data monitoring method as provided in the first aspect.
[0017] One or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages:
[0018] This application embodiment collects real-time port outbound and inbound traffic data of the monitored device, and introduces data processing and analysis algorithms based on artificial intelligence and deep learning in the backend to perform time-series collaborative analysis and dynamic interaction correlation of the port outbound and inbound traffic data of the monitored device, so as to perform real-time monitoring and anomaly detection of the monitored device's working status. This enables intelligent operation and maintenance monitoring of the monitored device on the cloud platform, allowing for timely detection of anomalies in the monitored device's working status, helping administrators to respond and handle issues more quickly, and improving system stability and security. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 A flowchart illustrating a data monitoring method provided in an embodiment of this application;
[0021] Figure 2 This is a schematic diagram of the structure of a data monitoring device provided in an embodiment of this application;
[0022] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0023] This application provides a data monitoring method that solves the technical problem in the prior art of being unable to effectively monitor data from different types of professional software and different computer terminals.
[0024] The technical solution of this application embodiment is to solve the above-mentioned technical problems, and the general idea is as follows:
[0025] A data monitoring method is applied to a cloud system that communicates with different types of monitored devices. The method includes: collecting operational status data of each monitored device according to a preset data collection method corresponding to each monitored device; the monitored devices include network devices, storage devices, server devices, middleware databases, big data clusters, and container clusters; collecting basic environmental data of each monitored device according to the corresponding environmental data collection method; integrating the basic environmental data and the operational status data corresponding to each monitored device to obtain monitoring fusion data, and visualizing the monitoring fusion data.
[0026] This application embodiment collects real-time port outbound and inbound traffic data of the monitored device, and introduces data processing and analysis algorithms based on artificial intelligence and deep learning in the backend to perform time-series collaborative analysis and dynamic interaction correlation of the port outbound and inbound traffic data of the monitored device, so as to perform real-time monitoring and anomaly detection of the monitored device's working status. This enables intelligent operation and maintenance monitoring of the monitored device on the cloud platform, allowing for timely detection of anomalies in the monitored device's working status, helping administrators to respond and handle issues more quickly, and improving system stability and security.
[0027] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.
[0028] First, it should be clarified that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0029] The exploration and development of unconventional natural gas refers to the exploration and development of natural gas reservoirs whose underground occurrence and accumulation patterns differ significantly from conventional natural gas reservoirs. During exploration and development, a large amount of data needs to be collected and professionally processed and analyzed. Related technologies mainly rely on cloud platforms for data resource sharing. However, because this data originates from different types of specialized software and different computer terminals, effective monitoring of this data is impossible. Therefore, how to effectively monitor data from different specialized software and different computer terminals is a pressing problem that needs to be solved.
[0030] To address the aforementioned problems, embodiments of this application provide, as follows: Figure 1 The data monitoring method shown is applied to a cloud system, which communicates with different types of monitored devices. The method includes steps S11-S13. The cloud system can adopt a B / S architecture, and the main operation can be performed based on a browser.
[0031] Step S11: Collect the operating status data of each monitored device according to the preset data collection method corresponding to each monitored device. Each monitored device includes network devices, storage devices, server devices, middleware databases, big data clusters and container clusters, as well as servers (hardware and operating system), storage, network devices, virtualization, cloud platform devices, and other devices, as well as IT infrastructure and virtualization devices such as network devices, server operating systems, databases, middleware, and storage, and computing servers, storage servers, network switches, routers, cloud desktops, security devices, host devices, business systems, etc. in the resource environment.
[0032] Step S12: Collect basic environmental data for each monitored device according to the environmental data collection method corresponding to each monitored device.
[0033] Step S13: Integrate the basic environmental data and the operating status data of each monitored device to obtain monitoring fusion data, and then visualize the monitoring fusion data.
[0034] Regarding step S11, the operating status data of each monitored device is collected according to the preset data collection method corresponding to each monitored device; each monitored device includes network devices, storage devices, server devices, middleware databases, big data clusters, and container clusters.
[0035] Due to the large number of monitored devices, to avoid traffic congestion at the data acquisition ports, distributed data acquisition servers can be deployed between the cloud system and the monitored devices. There can be multiple distributed data acquisition servers, each corresponding to at least one monitored device, with a one-to-one relationship between the distributed data acquisition servers and the monitored devices.
[0036] First, the distributed acquisition server collects the operating status data from the corresponding monitored devices, and then the cloud system collects the operating status data of each monitored device from the distributed acquisition server.
[0037] Specifically, it can control the distributed acquisition servers corresponding to each monitored device, and collect the operating status data of each monitored device according to the preset data acquisition method of each distributed acquisition server. It can also receive the operating status data of the monitored devices corresponding to each distributed acquisition server.
[0038] Specifically, when the computing power of the distributed acquisition server corresponding to each monitored device changes, the distributed acquisition server corresponding to each monitored device is controlled to readjust the monitoring correspondence between each distributed acquisition server and each monitored device according to the preset automatic adjustment mechanism.
[0039] Based on the adjusted monitoring correspondence, the distributed acquisition servers corresponding to each monitored device are controlled to collect the operating status data of their respective monitored devices according to the preset data acquisition methods of the monitored devices corresponding to each distributed acquisition server.
[0040] Changes in computing power can be caused by factors such as malfunctions or outages of distributed data acquisition servers. When computing power changes, the distributed data acquisition server may be unable to collect the operational status data of the monitored devices. To address this issue, a pre-set automatic adjustment mechanism can be used to readjust the monitoring relationships between the distributed data acquisition servers and the monitored devices, ensuring that the operational status data of the monitored devices can be collected normally and improving the monitoring efficiency.
[0041] The control and monitoring acquisition servers dynamically undertake their respective monitoring tasks based on the number of users and computing power. When the number of users and computing power change, they readjust their monitoring through an automatic adjustment mechanism.
[0042] The data acquisition and management function allows for customization of data acquisition methods for various devices, including network protocols, agents, APIs, and scripts, and controls probe consumption on the host through circuit breaker policies. The resource monitoring and management function provides the ability to monitor metrics data for the basic environment, network devices, storage devices, server devices, middleware databases, big data clusters, and container clusters, achieving integration of multiple metric data sources.
[0043] Furthermore, this application embodiment can manage all collectors through a data acquisition cluster, enabling operations such as stopping, starting, and removing the data acquisition server from the cluster. It can also manage all storage databases through the data acquisition cluster unit access mode, achieving health monitoring, task allocation, load balancing, and disaster recovery capabilities for the data acquisition modules, making the system more stable and reliable.
[0044] The monitoring data for server hardware includes: chassis connectivity, power status and operating status of the chassis and cage, chassis power status, chassis temperature, chassis power wattage, cage fan status, and cage power consumption. The monitoring data for storage devices includes: voltage, fan speed, power supply, and the CPU, controller, logic devices, disks, I / O modules, and connectivity of the storage device.
[0045] Operating system monitoring data includes: memory utilization, disk, CPU utilization, hard disk utilization, network card status, received and sent traffic and packet count, logs, Syslog, abnormal processes, number and size of directories and files, etc.
[0046] Monitoring data for network and security devices includes CPU and memory usage, interface status, port traffic, flow rate, packet loss rate, etc. It supports SNMPv1, v2, and v3 versions, and uses Syslog and SNMPTrap methods to collect event information from network devices.
[0047] Monitoring data for virtualization includes service success rate, average response time, CPU usage, memory usage, disk read / write performance, network reception rate, network transmission status, power status, and storage usage.
[0048] The monitoring data for the computer room includes monitoring access control status, reading access control entry and exit records, receiving smoke alarm information returned by smoke detection equipment, receiving alarm information from intelligent fire protection equipment, and monitoring equipment information such as air switches, power distribution cabinets, and fresh air units.
[0049] Regarding step S12, the basic environmental data of each monitored device is collected according to the environmental data collection method corresponding to each monitored device.
[0050] Basic environment data refers to the network environment that supports the normal operation of the monitored equipment, in addition to the equipment itself. Collecting basic environment data allows us to understand the status of the basic environment in which the monitored equipment resides, thereby improving the accuracy and efficiency of monitoring.
[0051] To ensure the high-speed and stable operation of cloud systems, servers, and monitored devices, cloud systems require monitoring and management of hardware devices and operating systems from multiple aspects. This includes collecting key operating parameters of server hardware such as CPU, memory, hard drive, and network card, as well as the running status of software and applications such as processes, services, and ports. System logs are categorized and scanned for queries. The system employs various methods including Agent, SNMP (V1, V2, V3), WMI, SSH, Telnet, IPMI, ILO, northbound interface, serial port, ODBC / JDBC, custom SQL, URL, WMI, and Java connections to uniformly collect configuration and indicator data from different vendors and types of servers, network devices, operating systems, storage, virtualization, middleware, databases, and web services. Mature modeling capabilities and indicator collection adaptation capabilities provide strong data source support for comprehensive operation and maintenance management of multiple types of devices.
[0052] Regarding step S13, the basic environmental data and the operating status data corresponding to each monitored device are integrated to obtain monitoring fusion data, and the monitoring fusion data is visualized.
[0053] Due to the large number of monitored devices, the corresponding operational status data and basic environmental data are extremely massive. Therefore, this operational status data and basic environmental data can be integrated to obtain fused monitoring data, which can then be visualized on display devices. The visualization page can display various types of operational data, including temperature, fan, disk, memory, bus, power supply, controller, virtualization layer, software, system, cloud desktop, and network metrics. Data integration supports different data sources such as Zabbix, MySQL, HTTP, File, Redis, MongoDB, Oracle, and PostgreSQL. It also supports data consumption for large-screen project management via API calls.
[0054] Data source integration can be achieved by classifying and filtering devices according to resource groups, device categories, or custom tags. Each device can belong to multiple resource groups, enabling on-demand execution of maintenance tasks and integrating the operational status data of devices belonging to the same category.
[0055] Visual representations and charts include, but are not limited to, line charts, pie charts, donut charts, column charts, bar charts, stacked charts, percentage stacked charts, waterfall charts, scatter plots, rose charts, funnel charts, query lists, top N charts, baseline charts, region maps, time graphs, flip charts, word clouds, dashboards, and rich text charts. Data views can be preset via SQL commands in the webpage for easy reading and display.
[0056] As can be seen, the embodiments of this application can perform unified monitoring and management of existing physical servers, switches, routers, cloud environments, and virtualization devices (such as VMware, XEN, Docker, KVM) of different types, models, and manufacturers, upgrading from distributed monitoring to centralized monitoring, and managing the entire IT system centrally and continuously from point to surface. This avoids problems such as difficulty in data correlation and sharing and lack of a unified operation and maintenance view in discrete monitoring, effectively realizing a global perspective for decision-making; at the same time, it can avoid the problems of high procurement costs, long construction time, high training costs, and large number of operation and maintenance personnel required caused by using multiple monitoring products.
[0057] To enhance the comprehensive management of IT operations and maintenance, this application's embodiments utilize existing open-source tools such as Zabbix, Grafana, and Prometheus, along with customized development capabilities, and integrate BI technology to develop a visualized operations and maintenance monitoring system. This system spans multiple virtualization layers and virtual machine cluster architectures, including physical servers, storage devices, data exchange devices, and virtualization layers such as XEN / VMware / Docker / KVM. It also features network topology self-discovery, network traffic detection, and network IP resource management. This system enables real-time monitoring, early warning, alarms, and multi-channel notifications of operational data from the cloud platform and other physical devices, virtualization layers, and network resources. This achieves unified management of IT system operations and maintenance, ensures controllable operation of cloud platform devices, improves the stability and reliability of the current IT platform, and adapts to the company's future IT development trends.
[0058] The cloud includes two communicating console servers. When the master console server fails, it controls the slave console server to collect the operating status data of each monitored device according to the preset data collection method corresponding to each monitored device.
[0059] The system supports dual-machine hot standby, with two main control console servers forming a highly efficient "master-slave" mode. The "master" and "slave" servers are linked in real time via "intelligent heartbeat" technology. If the master control server fails, the backup server immediately starts executing tasks. The technical architecture ensures that the system supports horizontal scaling, achieving a distributed architecture and providing architectural support for future terminal software management.
[0060] The system supports multi-machine disaster recovery backup, with multiple servers enabling mutual backup in case of server failure. If a server goes down, the monitoring tasks of these servers will be immediately redistributed to other normally operating servers, ensuring the continuity of monitoring.
[0061] In related technologies, due to the complexity of business operations, the alarm rules for each monitored device are configured differently, and alarm messages are stored in a scattered manner, making unified management difficult. Furthermore, in the IT business environment of these technologies, the monitored devices are interconnected. When some monitored devices malfunction, the cloud system often receives a large number of alarm messages, many of which are duplicates. Moreover, the alarm notification methods for different business operations vary. Therefore, when faced with a storm of alarm messages, it is impossible to analyze the existing alarm messages to identify those with critical value, thus hindering the rapid determination of the problem and its cause.
[0062] To address the aforementioned issues, in this embodiment of the application, when alarm information is included in the operational status data, the cloud system receives alarm information sent by each monitored device, resulting in an alarm information cluster. Then, according to the alarm object, alarm type, alarm level, and rule description attributes, the alarm information in the cluster is merged and compressed to obtain a simplified alarm cluster. Finally, alarm processing is performed on the alarm information within the simplified alarm cluster.
[0063] This application provides intelligent unified alarm management capabilities, enabling the unified management of multiple alarm sources from multiple monitored devices within a single system. Furthermore, it optimizes the alarm platform through an alarm suppression mechanism to handle large alarm storms. First, it compresses repetitive alarm messages, then merges similar alarms, and uses the alarm time sequence to determine the root cause. It supports various alarm notification methods, including but not limited to SMS, telephone, WeChat Work, email, sound, scripts, apps, work orders, and DingTalk, thereby improving overall operation and maintenance early warning and alarm capabilities.
[0064] Alarm templates can be pre-configured, and device information, monitoring point information, threshold settings, and fault time can be added to the alarm templates. Various alarm strategies can be set, including the following: sending an alarm after a certain number of consecutive occurrences of an event, sending an alarm after a set number of identical states within a set time period, stopping alarm sending after a certain number of consecutive occurrences of an event, and sending an alarm once when the monitoring point returns to normal after an alarm has been issued.
[0065] The alarm information processing of monitored devices includes message source integration, alarm suppression, event notification, and collaborative event handling. It enables configuration of alarm objects, alarm types, levels, and rule description attributes, and provides full lifecycle management of alarms, including discovery, viewing, receiving, dispatching, processing, improvement, and archiving. Notifications can be sent via multiple channels, including SMS, telephone, email, and instant messaging tools. It can access user operating system data and multiple alarm data sources, generating alarm events through alarm compression, merging, and push notifications to relevant personnel for processing.
[0066] As can be seen, the embodiments of this application monitor and manage IT infrastructure and virtualization, including network devices, server operating systems, databases, middleware, and storage. Through various technical protocols and methods such as SNMP, IPMI, PING, HTTP, SSH, Telnet, WMI, and Agent, it achieves uninterrupted monitoring of computing servers, storage servers, network switches, routers, cloud desktops, security devices, host devices, and business systems within the resource environment. It focuses on the health of key indicators, promptly predicts potential problems, and automatically alerts relevant personnel upon detecting anomalies via SMS and email. By analyzing the monitoring data, it displays the real-time operating status and business health of devices, networks, and system software through charts and topology diagrams, providing operations and maintenance personnel with intuitive and efficient decision-making support and improving monitoring efficiency.
[0067] In response to the current situation where alarms are not timely, have low accuracy, and have high usage barriers, a unified monitoring and alarm system is built to unify the access and processing of alarm messages and data indicators from various monitoring systems. This system supports the filtering, notification, response, handling, classification, tracking, and multi-dimensional analysis of alarm events, enabling global control over the entire lifecycle of problem events and improving monitoring efficiency.
[0068] This system provides a visual representation of various operational data, including temperature, fan speed, disk usage, memory, bus speed, power supply, controller speed, virtualization layer speed, software performance, system performance, cloud desktop performance, and network speed. It supports data access via static mock data, HTTP interfaces, and external databases, and allows for flexible modification of data structures. Based on open-source templates and custom development methods, it generates corresponding icon content and supports intranet deployment. The UI, color scheme, and resolution for large-screen displays are all configurable.
[0069] The visualization in this application relies on the VUE framework for development. Data integration supports different data sources such as Zabbix, MySQL, HTTP, File, Redis, MongoDB, Oracle, and PostgreSQL. It also supports data consumption for the management and use of large-screen projects via API calls.
[0070] The solution provided in this application provides integrated monitoring of IT equipment and power environment in data centers. It supports comprehensive and in-depth monitoring of IT resources such as network equipment, servers, storage, and virtual resources from various manufacturers, as well as data center environmental monitoring such as UPS, air conditioning, temperature and humidity control, smoke detection, and water immersion monitoring, ensuring the continuous and stable operation of the network and IT systems. A customized large-screen display can also be provided to show the operational status, health status, and alarm information of infrastructure. This allows maintenance personnel to monitor the overall system operation in real time through the dynamic display, promptly address any anomalies, ensure the safe and stable operation of the system, and improve monitoring efficiency.
[0071] In addition, embodiments of this application also provide testing and prevention of the monitoring effect of the monitoring method, which may specifically include business / requirement driven testing, product quality risk driven testing, function driven testing, and structure driven testing.
[0072] Business / requirement-driven testing: Starting from the actual business needs of users, this involves analyzing test objects such as business objectives, business processes, user roles, business rules, business data, and business development. For these objects, the testing scope, methods, and strategies are determined. The adequacy of the testing is measured using business processes and data. Requirement-based verification methods and user scenario-based testing methods are used to determine whether the software system fully meets business needs.
[0073] Product quality risk-driven testing: First, internal quality is assessed to expose quality risks, and then it is gradually expanded to assess external system quality and user usage quality, continuously revealing and providing feedback on the main product quality risks.
[0074] Function-driven testing: Starting from the system's functional characteristics and based on the requirements specification, each function is verified to determine whether it operates normally and whether it matches the preset effect. Specifically, it can be divided into sub-functions and sub-functions of sub-functions, forming a list of functional points, and test cases are designed and executed for each functional point.
[0075] Structure-driven testing includes structured testing and white-box testing. Testing is driven by the program structure, involving structural analysis to progressively cover each part of the program and its relationships, such as component-based testing, interface-based testing, or API-based testing. It also involves testing based on code structure, including line coverage, branch coverage, and basic path coverage. Structure-driven testing offers more objective sufficiency measurements, especially based on code coverage analysis. This application's embodiments utilize relevant tools for automated testing.
[0076] Furthermore, after collecting the operational status data of each monitored device and the basic environmental data of each monitored device, this embodiment of the application determines whether there is an anomaly in any target monitored device among the monitored devices based on the first time series of port outflow traffic and the second time series of port inflow traffic in the operational status data corresponding to the target monitored device, specifically including steps S21-S24.
[0077] Step S21: Extract time-series pattern features from the first time series and the second time series, and perform semantic enhancement processing on the first time series and the second time series to obtain the first enhanced port feature vector and the second enhanced port feature vector.
[0078] Step S22: Determine the mean feature vector of the first enhanced port feature vector and the second enhanced port feature vector respectively, and determine the first associated feature vector and the second associated feature vector.
[0079] Step S23: Determine the port ingress / egress interaction vector based on the first associated feature vector and the second associated feature vector;
[0080] Step S24: Determine whether there is any abnormality in the target monitored device based on the port inbound and outbound interaction vectors.
[0081] The first time series of port outflow and the second time series of port inflow are used to extract features from a traffic time series pattern feature extractor based on a Bi-LSTM model. This extracts local temporal contextual features of port outflow and inflow, resulting in the first and second local time series correlation sequences. It should be understood that the Bi-LSTM model is a deep learning model suitable for processing time series data, effectively capturing long-term dependencies in time series data. In monitoring systems, changes in port outflow and inflow exhibit significant temporal characteristics; therefore, using Bi-LSTM allows for better analysis and understanding of this time series data. In particular, the unique feature of the Bi-LSTM model is its ability to process both past and future information simultaneously. This means the model can not only consider the impact of historical traffic data on the current state but also utilize future traffic data to enhance the understanding of the current state. This provides a more accurate and reliable foundation for subsequent anomaly detection of monitored equipment.
[0082] The first time series of port outflow and the second time series of port inflow are subjected to a traffic time series pattern feature extractor based on the Bi-LSTM model to obtain a first local time series correlation sequence and a second local time series correlation sequence. The first local time series correlation sequence and the second local time series correlation sequence are subjected to saliency-global contextual semantic enhancement processing to obtain a first enhanced port feature vector and a second enhanced port feature vector.
[0083] Then, considering that the first and second local time-series correlation sequences respectively contain local time-series contextual feature information about port outbound and port inbound traffic in the time dimension, and that these local time-series features have a global correlation across the entire time domain, and considering that these local time-series features of port outbound and port inbound traffic are not equally important during the anomaly detection process of the monitored device's operating status, the technical solution of this application further performs saliency-global contextual semantic enhancement processing on the first and second local time-series correlation sequences to obtain a first enhanced port feature vector and a second enhanced port feature vector. Through saliency-global contextual semantic enhancement processing, the global correlation relationship between the various local time-series features of port outbound and port inbound traffic in the time dimension and the importance features related to subsequent operating status monitoring tasks can be automatically learned and captured, and these important semantics are saliency-highlighted to obtain a more comprehensive description of the operating status of the monitored device, thereby achieving more accurate operation and maintenance monitoring of the monitored device on the cloud platform.
[0084] The mean and maximum values of the local temporal correlation feature vectors of each port outflow in the first local temporal correlation sequence are calculated respectively to obtain the first local global semantic vector composed of the mean values of multiple vectors and the first local salient semantic vector composed of the maximum values of multiple vectors. Contextual feature capture is performed on the first local global semantic vector to obtain the first local global latent vector. Contextual feature capture is also performed on the first local salient semantic vector to obtain the first local salient latent vector. The first local global latent vector and the first local salient latent vector are fused to obtain the first local global salient vector. After obtaining the first local global salient vector, nonlinear activation processing is performed on the first local global salient vector to obtain the first local global salient weight vector. Using the first local global salient weight vector as the weight, the first local temporal correlation sequence is multiplied by position and then added to the first local temporal correlation sequence to obtain the second enhanced port feature vector.
[0085] Furthermore, the first local global semantic vector is encoden by one-dimensional convolution and then processed by the ReLU activation function to obtain the activated expression of the first local global semantic vector; the activated expression of the first local global semantic vector is then encoden by one-dimensional convolution again and multiplied by the first trainable linear transformation weight matrix to obtain the first local global latent vector.
[0086] Furthermore, the first local salient semantic vector is encoden by one-dimensional convolution and then processed by the ReLU activation function to obtain the activated expression of the first local salient semantic vector; the activated expression of the first local salient semantic vector is then encoden by one-dimensional convolution again and multiplied by the second trainable linear transformation weight matrix to obtain the first local salient latent vector.
[0087] Furthermore, the first local global latent vector and the first local salient latent vector are fused to obtain the first local global salient vector; the first local global salient vector is activated by the hyperbolic tangent function and the sigmoid function respectively to obtain the first local global salient tangent vector and the first local global salient S vector; the first local global salient tangent vector and the first local global salient S vector are then multiplied by position to obtain the first local global salient weight vector.
[0088] Specifically, the first local temporal association sequence is subjected to contextual semantic enhancement processing based on saliency-globality, and is processed with the following saliency enhancement formula to obtain the first enhanced port feature vector;
[0089] The formula for saliency enhancement is as follows:
[0090] V′=Avg(V)
[0091] V" = Max(V)
[0092]
[0093] V c =V⊙(tanh(C1+C2)⊙σ(C1+C2))+V
[0094] Where V represents the local temporal correlation feature vector of each port outflow in the first local temporal correlation sequence, Avg(·) denotes the global mean operation in the vector, Max(·) denotes the maximum value operation in the vector, V′ is the semantic mean representation vector of the local temporal correlation of port outflow, V" is the semantic maximum value representation vector of the local temporal correlation of port outflow, Conv(·) denotes one-dimensional convolutional coding, and ReLU(·) denotes the ReLU activation function. and Let C1 and C2 represent the trainable weight matrices used for linear transformation, respectively. C1 and C2 are the feature vectors representing the mean and maximum values of the local temporal association hidden semantics of port outflow, respectively. σ(·) is the Sigmoid function, tanh(·) is the tanh function, and ⊙ represents the positional dot product. c This is the local time-series correlation feature vector of the outflow of each enhanced port in the first enhanced port feature vector.
[0095] Next, the positional mean vectors of the first and second reinforced port feature vectors are calculated respectively to obtain the first and second associated feature vectors. By performing mean calculations based on multiple vectors in the sequence for each position in the first and second reinforced port feature vectors, the temporal information in the sequence can be aggregated into a single feature vector. This helps to capture the overall behavioral patterns of port outflow and inflow within the entire time window, helps to capture and understand the overall trend and periodicity of the sequence data, and helps the system to more comprehensively understand the dynamic changes and trends of port flow data. This provides a more effective feature representation and analysis foundation for subsequent feature dynamic interaction and monitoring result generation.
[0096] Furthermore, since the first and second associated feature vectors respectively contain full-time temporal correlation representation features of port outflow traffic and port inflow traffic, there are complex dynamic correlations and interactions between these two features. Therefore, in order to better capture the dynamic relationship and temporal interaction features between port outflow and inflow traffic, and to provide a more comprehensive feature representation for subsequent anomaly detection and monitoring result generation, the technical solution of this application further performs feature dynamic interaction processing based on prior distribution on the first and second associated feature vectors to obtain port outflow-inflow interaction vectors. Through feature dynamic interaction processing based on prior distribution, the temporal features of port outflow and inflow traffic can be used as prior information to guide the learning and capture of dynamic interaction feature information between the temporal features of port outflow and inflow traffic, so as to more comprehensively reflect the temporal changes in the relationship between the two, thereby improving the system's ability to detect abnormalities in the device's operating status and helping to more accurately determine whether there are abnormalities in the device's operating status.
[0097] Calculate the position-by-position division between the first associated feature vector and the second associated feature vector to obtain the position-by-position feature vector for entry and exit; calculate the first prior factor and the second prior factor of the position-by-position feature vector for entry and exit respectively to obtain the first prior feature vector for entry and exit and the second prior feature vector for entry and exit; calculate the position-by-position summation between the first prior feature vector for entry and exit and the second prior feature vector for entry and exit to obtain the port entry and exit interaction vector.
[0098] Further, the positional feature vectors of entry and exit are multiplied position by position by a first predetermined weight hyperparameter to obtain a first weighted entry and exit response feature vector; the positional feature values in the first weighted entry and exit response feature vector are used as exponents of the natural constant to calculate the positional exponential function value with the natural constant as the base, to obtain a first weighted entry and exit support feature vector; the first entry and exit support feature vector is multiplied by a first Gaussian random number function value to obtain a first entry and exit prior feature vector; the positional feature vectors of entry and exit are multiplied position by position by a second predetermined weight hyperparameter to obtain a second weighted modulated entry and exit positional feature vector; the positional feature values in the second weighted modulated entry and exit positional feature vector are used as exponents of the natural constant to calculate the positional exponential function value with the natural constant as the base, to obtain a second weighted entry and exit support feature vector; the second weighted entry and exit support feature vector is multiplied by a second Gaussian random number function value to obtain a second entry and exit prior feature vector; wherein, the first Gaussian random number function value and the second Gaussian random number function value are both generated by a Gaussian random number function with a mean of 0 and a variance of 1.
[0099] Specifically, the first associated feature vector and the second associated feature vector are subjected to dynamic feature interaction processing based on prior distribution, and the following dynamic feature interaction formula is used to obtain the port inlet and outlet interaction vector.
[0100] The feature dynamic interaction formula is as follows:
[0101]
[0102] Where v1 is the first associated feature vector, v2 is the second associated feature vector, and p1 and p2 are predetermined weight hyperparameters. These are the positional feature vectors for entry and exit. f1 and f2 are hyperparameters of the Gaussian distribution function, which generates Gaussian random numbers with a mean of 0 and a variance of 1. For vector addition, V c This is the port inbound / outbound interaction vector.
[0103] Furthermore, the port inbound and outbound interaction vectors are processed by a classifier-based operation and maintenance monitoring result generator to obtain the operation and maintenance monitoring results. These results indicate whether there are any abnormalities in the working status of the monitored devices. In other words, the dynamic interaction characteristics between the time-series features of port inbound and outbound traffic are used for classification processing to perform real-time monitoring and anomaly detection of the monitored devices' working status. This enables intelligent operation and maintenance monitoring of the monitored devices on the cloud platform, allowing for timely detection of anomalies in the monitored devices' working status, helping administrators respond and handle issues more quickly, and improving system stability and security.
[0104] The port inbound and outbound interaction vectors are processed by a classifier-based operation and maintenance monitoring result generator to obtain operation and maintenance monitoring results. These results are used to indicate whether there are any abnormalities in the working status of the monitored equipment.
[0105] Furthermore, a traffic temporal pattern feature extractor based on the Bi-LSTM model, a contextual semantic enhancement processing based on saliency-globality, a feature dynamic interaction processing based on prior distribution, and a maintenance monitoring result generator based on a classifier are trained. Based on the first time series of outbound traffic and the second time series of inbound traffic from the training port of the monitored equipment, as well as the true value indicating whether the working status of the monitored equipment is abnormal, temporal pattern features are extracted from the first time series of outbound traffic and the second time series of inbound traffic from the training port, and semantic enhancement processing is performed on the first and second time series to obtain the first enhanced port feature vector and the second enhanced port feature vector. The first enhanced port feature vector and the second enhanced port feature vector are then calculated respectively. The position-mean vector of the second enhanced port feature vector is trained to obtain the first and second associated feature vectors. The first and second associated feature vectors are then subjected to dynamic feature interaction processing based on prior distribution to obtain the port in-and-out interaction vector. The port in-and-out interaction vector is then processed by a classifier-based operation and maintenance monitoring result generator to obtain the operation and maintenance monitoring results. The cross-entropy loss function value between the operation and maintenance monitoring results and the true value is calculated to obtain the classification loss function value. Based on the classification loss function value, the traffic time-series pattern feature extractor based on the Bi-LSTM model, the context semantic enhancement processing based on saliency-globality, the dynamic feature interaction processing based on prior distribution, and the classifier-based operation and maintenance monitoring result generator are trained.
[0106] Preferably, the input and output interaction vectors of the training ports are used to generate operation and maintenance monitoring results based on a classifier, resulting in operation and maintenance monitoring results including:
[0107] Multiply each feature value of the training port inflow-outflow interaction vector by the length of the training port inflow-outflow interaction vector, subtract the first norm of the training port inflow-outflow interaction vector, and calculate the square root of the absolute value of the subtraction result to obtain the first port inflow-outflow time-series dynamic interaction intermediate vector.
[0108] Multiply each feature value of the training port inflow-outflow interaction vector by the square root of the length of the training port inflow-outflow interaction vector, subtract the L2 norm of the training port inflow-outflow interaction vector, and calculate the square root of the absolute value of the subtraction result to obtain the second port inflow-outflow time-series dynamic interaction intermediate vector.
[0109] Calculate the base-2 logarithm of each feature value of the first port inflow-port outflow time-series dynamic interaction intermediate vector and the second port inflow-port outflow time-series dynamic interaction intermediate vector to obtain the first port inflow-port outflow time-series dynamic interaction intermediate information vector and the second port inflow-port outflow time-series dynamic interaction intermediate information vector.
[0110] The corrected port ingress / egress interaction vector is obtained by calculating the weighted sum of the intermediate information vectors of the time-series dynamic interaction between the first port ingress and the second port ingress / egress traffic; and
[0111] The corrected port inbound and outbound interaction vectors are processed by a classifier-based operation and maintenance monitoring result generator to obtain the operation and maintenance monitoring results.
[0112] Here, for the local temporal short-to-long-range bidirectional temporal correlation features of the training port outflow and inflow, expressed by the first and second correlation feature vectors respectively, after performing local temporal saliency-global context enhancement and local temporal mean regression in the global temporal domain, due to the difference in the distribution of feature dynamic interaction weights based on the prior distribution caused by the difference in local-global distribution in the source temporal domain, the training port outflow-inflow interaction vector will also have the problem of insufficient aggregation of interaction temporal feature distribution, thus affecting its classification convergence efficiency through the classifier, that is, affecting the efficiency of classification training and the accuracy of classification results.
[0113] Based on this, the applicant of this application uses the training port inbound / outbound interaction vector as a feature set, where the temporal features are represented by the feature values of the training port inbound / outbound interaction vector. To dynamically aggregate the overall temporal feature set composed of different temporal features of the training port inbound / outbound interaction vector without ignoring individual temporal feature changes, the individual features of the training port inbound / outbound interaction vector and their aggregated scale are represented as the full amplitude and half amplitude of the features. Furthermore, the low-rank negative correlation of different dimensions of the overall temporal features of the training port inbound / outbound interaction vector feature set is used as the phase and scaled to dynamically adjust the changing relationships of the temporal feature content of the training port inbound / outbound interaction vector. This improves the aggregation of the overall temporal feature information of the training port inbound / outbound interaction vector feature set, and improves the classification convergence efficiency of the training port inbound / outbound interaction vector through the classifier, i.e., improving the efficiency of classification training and the accuracy of classification results. In this way, the working status of the monitored devices on the cloud platform can be more accurately monitored in real time and anomaly detected, thereby realizing intelligent operation and maintenance monitoring of the monitored devices on the cloud platform. This allows for timely detection of anomalies in the working status of the monitored devices, helping administrators to respond and handle issues more quickly, and improving the stability and security of the system.
[0114] In summary, the above solution involves real-time monitoring and collection of port outbound and inbound traffic data from the monitored devices. Backend data processing and analysis algorithms based on artificial intelligence and deep learning are then introduced to perform time-series collaborative analysis and dynamic interaction correlation of this data. This enables real-time monitoring and anomaly detection of the monitored devices' operational status. This achieves intelligent operation and maintenance monitoring of the monitored devices on the cloud platform, allowing for timely detection of anomalies in their operational status, helping administrators respond and handle issues more quickly, and improving system stability and security.
[0115] (1) Comprehensive network topology management, providing fast and accurate topology discovery. The topology search function can be switched to run in the background so that the operation of other functions is not affected during the topology search. Topology management should support functions such as topology discovery, device overview, link overview, device location, topology navigation, topology display, topology thumbnail, large screen display, and topology customization.
[0116] (2) Network and security device monitoring: Supports automatic discovery modes such as full network discovery, extended discovery, and network segment discovery. It supports both automatic discovery and manual addition for monitoring network devices and can automatically draw network topology diagrams. It can collect and provide early warnings for performance indicators of routers, switches, security devices, load balancers, and other devices that meet SNMPv1 / v2 / v3 standards. It uses ICMP and SNMPGET for device detection; the system has built-in templates supporting indicators such as CPU utilization, memory utilization, port status, port inbound and outbound traffic, inbound and outbound error frame rates, broadcast inbound and outbound frame rates, ping latency, and packet loss. It supports monitoring of network device hardware information, such as the status of fans, power supplies, and temperatures.
[0117] (3) Configuration file management: Supports automatic acquisition and backup of configuration information of mainstream online network devices via SSH / Telnet; supports importing device resource libraries. Supports scheduled automated task management for devices; supports viewing detailed contents of configuration files; supports comparison between configuration files, marking and displaying changes. Supports management of backup scripts by device model.
[0118] (4) Storage Monitoring: The system must support monitoring and management of NAS storage, providing monitoring of storage devices from mainstream manufacturers. This includes: IBM DS disk arrays, IBM NetAPP disk arrays, EMCVNX / UNX / CX / NS disk arrays, HPMSA disk arrays, HP3PAR disk arrays, Hitachi HDS disk arrays, Huawei OceanStor disk arrays, Inspur AS disk arrays, Synology NAS, Fujitsu disk arrays, Hikvision DS disk arrays, and other storage devices. The system automatically monitors the operating status of disks, storage volumes, and controllers within the disk array, providing timely information on performance status, operating status, list of faulty components, IO interval time, cache hit count, cache hit time, cache hit rate, number of bytes transferred, transfer rate, and cumulative idle time, displaying the real-time performance status of the storage devices.
[0119] (5) Alarm Event and Response Management: Alarm management should include alarm views and alarm policy settings. It provides network alarm monitoring, promptly notifying users via SMS, email, etc., upon the occurrence of faults, and offering alarm analysis and statistical reports to provide proactive fault resolution methods for maintenance personnel. The monitoring center platform software can act as a log information receiving server for IT resources, supporting log collection protocols including Syslog, SNMPTrap, and agents. It can continuously collect event logs or system logs related to operating status, user behavior, security events, and hardware alarms from distributed Windows, Linux, AIX hosts, application systems, networks, and security devices. Log receiving filtering conditions can be set, storing logs that conform to the rules in the database. Alarm policies can be customized, and regular expressions can be configured for processing. Alarms can be generated, and log information can be converted into system alarms based on the content of the alarm message. Multiple alarm methods are provided, supporting email, SMS, WeChat Work, DingTalk, and other alarm methods.
[0120] (6) Equipment inspection management: Inspection plans support both manual and automatic execution modes. The chapters and inspection items of the inspection report can be customized. The inspection type supports system self-inspection and manual inspection. System self-inspection allows batch selection of inspection resource types and monitoring indicators. Inspection reports support PDF and Word formats.
[0121] (7) Reporting system: The system should be able to generate categorized reports or unified reports for monitored resources, including network devices, application services, storage, and virtual machines, presenting all resource data managed in the system in a single report. It should also be customizable to display various resources as needed. The system should provide performance reports, alarm statistics reports, TOPN reports, availability reports, trend reports, analysis reports, and comprehensive reports. The report management system should have query functionality, allowing for customized report display based on resources and indicators, and customizable views.
[0122] (8) System security: The portal system has robust security mechanisms, including terminal IP address access filtering, forced password modification, and prevention of brute-force attacks. Furthermore, all user passwords and data information involved in the system are stored in encrypted form. The portal records operation logs for all logins and login operations, as well as login successes and failures. All add, delete, and modify operations are also recorded in operation logs, which provide a query function.
[0123] (9) Third-party system interfaces allow different data users to access the operation and maintenance management platform through standard data interfaces, and allow different data providers to join the operation and maintenance management platform through standard data interfaces. The operation and maintenance management platform effectively integrates with third-party systems and security management platforms to achieve functions such as unified alarm display and unified authentication login integration; it can effectively integrate and unify risk alarms of various security devices in the user network. The interfaces comply with ITSS specifications.
[0124] (10) Data integration: Interface integration with Citrix Hypervisor 8.2 + XenCenter 8.2 virtualization platform and VMware 8.0.1 (vSphere + vSAN + vCenter) virtualization platform to collect specified metrics. Data integration with DELL EMC PowerStore 1000T direct-attached storage system, ES-5000D distributed storage system, and S3 object storage system to collect specified metrics.
[0125] The security of the integrated business operation and maintenance management platform is a fundamental factor in ensuring the normal operation of user management. Therefore, the security of the management system was fully considered during the product architecture design, and specific measures included:
[0126] 1. Configure system firewall policies to allow only IP addresses and ports necessary for the business system, and prohibit access to non-specified IP addresses and ports.
[0127] 2. The system supports HTTPS access.
[0128] 3. Limit the number of failed user login attempts, enforce password changes, and prevent brute-force attacks.
[0129] 4. A strict access control mechanism is adopted, with detailed permissions and management scope for each user.
[0130] 5. Log user operations in detail to meet security audit requirements.
[0131] 6. All user passwords and connection information of managed devices in the system are stored using DES encryption to prevent unauthorized access to device connection information.
[0132] In the integrated operation and maintenance monitoring method for monitored devices on the cloud platform, the software architecture is designed so that all structured data on the platform is permanently stored in a MySQL database. The basic equipment monitoring platform is divided into three layers: a data acquisition layer, a data analysis and processing layer, and a data display layer. The platform adopts a modular design with loose coupling between modules. New modules can be directly connected to the platform, and modules communicate with each other through interfaces, message queues, and other methods. The data acquisition layer is the foundation of the entire management platform, responsible for collecting the data required for platform operation. The data acquisition layer obtains the necessary indicator information from the managed devices through various network protocols, including SNMP, SSH, IPMI, PING, JDBC, JMX, and SMI-S. The collected data is cached for parsing and processing, and then stored in the database for analysis and display by the upper-layer platform. The platform has a built-in extensible resource capability library model. For manufacturers, models, and indicators that do not meet the requirements, the system can be configured without secondary development. It supports custom expansion of monitoring indicators through SNMP, JDBC, JMX, and other methods. The operation and maintenance monitoring method can be deployed on the same or different hosts according to the actual situation of the customer's IT environment through MHS (Information Processing) service and MCS (Information Collection) service. At the same time, depending on the scale of the customer's managed objects, one or more MCSs can be used for management capacity planning, realizing two different deployment methods, namely centralized or distributed, and achieving flexible management of IT resources.
[0133] In summary, this application embodiment collects port outbound and inbound traffic data of the monitored device in real time, and introduces data processing and analysis algorithms based on artificial intelligence and deep learning in the backend to perform time-series collaborative analysis and dynamic interactive correlation of the port outbound and inbound traffic data of the monitored device, so as to perform real-time monitoring and anomaly detection of the monitored device's working status. This enables intelligent operation and maintenance monitoring of the monitored device on the cloud platform, allowing for timely detection of anomalies in the monitored device's working status, helping administrators to respond and handle issues more quickly, and improving system stability and security.
[0134] Based on the same inventive concept, the embodiments of this application provide, as follows: Figure 2 The data monitoring device shown is applied to a cloud system, which communicates with different types of monitored devices. The device includes:
[0135] The data acquisition module 21 is used to collect the operating status data of each monitored device according to the preset data acquisition method corresponding to each monitored device; each monitored device includes network devices, storage devices, server devices, middleware databases, big data clusters and container clusters.
[0136] Data acquisition module 21 is used to collect basic environmental data of each monitored device according to the environmental data acquisition method corresponding to each monitored device;
[0137] The fusion display module 22 is used to integrate the basic environmental data and the operating status data of each monitored device to obtain the monitoring fusion data, and to visualize the monitoring fusion data.
[0138] Furthermore, the data acquisition module 21 is used for:
[0139] Control the distributed acquisition servers corresponding to each monitored device, and collect the operating status data of each monitored device according to the preset data acquisition method of each monitored device.
[0140] It receives the operating status data of the monitored devices from the distributed acquisition servers corresponding to each monitored device.
[0141] Furthermore, the data acquisition module 21 is used for:
[0142] When the computing power of the distributed acquisition server corresponding to each monitored device changes, the distributed acquisition server corresponding to each monitored device is controlled to readjust the monitoring correspondence between each distributed acquisition server and each monitored device according to the preset automatic adjustment mechanism.
[0143] Based on the adjusted monitoring correspondence, the distributed acquisition servers corresponding to each monitored device are controlled to collect the operating status data of their respective monitored devices according to the preset data acquisition methods of the monitored devices corresponding to each distributed acquisition server.
[0144] Furthermore, the cloud includes two communicating console servers and a data acquisition module for:
[0145] When the master console server of the two console servers fails, the slave console server of the two console servers is controlled to perform the steps of collecting the operating status data of each monitored device according to the preset data collection method corresponding to each monitored device.
[0146] Furthermore, the alarm information module is used for:
[0147] Receive alarm information sent by each monitored device to obtain an alarm information cluster;
[0148] According to the alarm object, alarm type, alarm level and rule description attributes, the alarm information in the alarm information cluster is merged and compressed to obtain a simplified alarm cluster.
[0149] Perform alarm processing on alarm information existing in the streamlined alarm cluster.
[0150] Furthermore, the anomaly detection module is used for:
[0151] After collecting operational status data and basic environmental data of each monitored device, for any target monitored device, the system determines whether there is an anomaly based on the first time series of port outflow traffic and the second time series of port inflow traffic in the operational status data of the target monitored device.
[0152] Furthermore, the anomaly detection module is used for:
[0153] Temporal pattern features are extracted from the first time series and the second time series, and semantic enhancement processing is performed on the first time series and the second time series to obtain the first enhanced port feature vector and the second enhanced port feature vector.
[0154] Determine the mean feature vector of the first enhanced port feature vector and the second enhanced port feature vector respectively, and determine the first associated feature vector and the second associated feature vector;
[0155] The port ingress / egress interaction vectors are determined based on the first and second associated feature vectors.
[0156] Based on the port inbound and outbound interaction vectors, determine whether there are any anomalies in the target monitored device.
[0157] Based on the same inventive concept, the embodiments of this application provide, as follows: Figure 3 An electronic device shown includes:
[0158] Processor 31;
[0159] Memory 32 is used to store executable instructions of processor 31;
[0160] The processor 31 is configured to execute a data monitoring method as described above.
[0161] Based on the same inventive concept, embodiments of this application provide a non-transitory computer-readable storage medium that, when the instructions in the storage medium are executed by the processor 31 of an electronic device, enables the electronic device to perform a data monitoring method as described above.
[0162] Since the electronic device described in this embodiment is an electronic device used to implement the information processing method in the embodiments of this application, those skilled in the art can understand the specific implementation methods and various variations of the electronic device in this embodiment based on the information processing method described in the embodiments of this application. Therefore, how the electronic device implements the method in the embodiments of this application will not be described in detail here. Any electronic device used by those skilled in the art to implement the information processing method in the embodiments of this application falls within the scope of protection of this application.
[0163] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0164] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0165] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0166] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1The steps of the function specified in one or more boxes.
[0167] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.
[0168] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. A data monitoring method, characterized in that, Applied to a cloud system that communicates with different types of monitored devices, the method includes: According to the preset data acquisition method corresponding to each of the monitored devices, the operating status data of each monitored device is collected; each monitored device includes network devices, storage devices, server devices, middleware databases, big data clusters, and container clusters; According to the environmental data acquisition method corresponding to each of the monitored devices, the basic environmental data of each of the monitored devices are collected; The basic environmental data and the operating status data corresponding to each of the monitored devices are integrated to obtain monitoring fusion data, and the monitoring fusion data is visualized.
2. The method as described in claim 1, characterized in that, The step of collecting operational status data of each monitored device according to a preset data collection method corresponding to each monitored device includes: Control the distributed acquisition servers corresponding to each of the monitored devices, and collect the operating status data of the monitored devices according to the preset data acquisition method corresponding to each of the monitored devices; The system receives the operating status data of the monitored devices from the distributed acquisition servers corresponding to each of the monitored devices.
3. The method as described in claim 2, characterized in that, The control of each of the monitored devices' corresponding distributed acquisition servers, according to the preset data acquisition method for each monitored device, involves collecting the operating status data of each monitored device, including: When the computing power of the distributed acquisition server corresponding to each of the monitored devices changes, the distributed acquisition server corresponding to each of the monitored devices is controlled to readjust the monitoring correspondence between each of the distributed acquisition servers and each of the monitored devices according to a preset automatic adjustment mechanism. Based on the adjusted monitoring correspondence, the distributed acquisition servers corresponding to each monitored device are controlled to collect the operating status data of their respective monitored devices according to the preset data acquisition method of the monitored devices corresponding to each distributed acquisition server.
4. The method as described in claim 1, characterized in that, The cloud includes two console servers that communicate with each other, and the method further includes: When the master console server of the two console servers malfunctions, the slave console server of the two console servers is controlled to perform the step of collecting the operating status data of each of the monitored devices according to the preset data collection method corresponding to each of the monitored devices.
5. The method as described in claim 1, characterized in that, The method further includes: Receive alarm information sent by each of the monitored devices to obtain an alarm information cluster; According to the alarm object, alarm type, alarm level and rule description attributes, the alarm information in the alarm information cluster is merged and compressed to obtain a simplified alarm cluster. The alarm information existing in the simplified alarm cluster is processed.
6. The method as described in claim 1, characterized in that, After collecting operational status data and basic environmental data of each of the monitored devices, the method further includes: For any one of the monitored devices, determine whether the target monitored device has any abnormalities based on the first time series of port outflow and the second time series of port inflow in the operating status data corresponding to the target monitored device.
7. The method as described in claim 6, characterized in that, The step of determining whether the target monitored device has any abnormalities based on the first time series of port outbound traffic and the second time series of port inbound traffic in the operating status data corresponding to the target monitored device includes: Temporal pattern features are extracted from the first time series and the second time series, and semantic enhancement processing is performed on the first time series and the second time series to obtain a first enhanced port feature vector and a second enhanced port feature vector. The mean feature vectors of the first enhanced port feature vector and the second enhanced port feature vector are determined respectively, and the first associated feature vector and the second associated feature vector are determined. Based on the first associated feature vector and the second associated feature vector, determine the port ingress / egress interaction vector; Based on the port inbound and outbound interaction vectors, determine whether the target monitored device has any abnormalities.
8. A data monitoring device, characterized in that, An apparatus for use in a cloud system that communicates with different types of monitored devices, the apparatus comprising: The data acquisition module is used to collect the operating status data of each of the monitored devices according to the preset data acquisition methods corresponding to each of the monitored devices; each of the monitored devices includes network devices, storage devices, server devices, middleware databases, big data clusters, and container clusters. The data acquisition module is used to collect basic environmental data of each of the monitored devices according to the environmental data acquisition method corresponding to each of the monitored devices; The fusion display module is used to integrate the basic environmental data and the operating status data corresponding to each of the monitored devices to obtain fused monitoring data, and to visualize the fused monitoring data.
9. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute a data monitoring method as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium, wherein instructions in the storage medium, when executed by a processor of an electronic device, enable the electronic device to perform a data monitoring method as described in any one of claims 1 to 7.