Distributed Operation and Maintenance Intelligent Management Method and System Based on Cloud Computing and Internet of Things
By generating fault logic diagram encoding vectors and permeability node vectors in a cloud computing environment, combined with deep learning network prediction fault root cause description labels, the problem of time-consuming and laborious fault location in traditional operation and maintenance methods is solved, and the efficient and accurate identification and management of fault activities is achieved.
Patent Information
- Application Number
- CN202411788421.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-06
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2044-12-06
AI Technical Summary
Traditional operation and maintenance methods are difficult to effectively deal with fault location and management in the interactive event flow of IoT devices in cloud computing environments, especially the accurate positioning of the fault chain is time-consuming and labor-intensive and error-prone.
By obtaining the inference fault knowledge chain of potential fault activities, a fault logic diagram encoding vector and fault penetration node vector are generated, and a target stacking vector is generated by combining the first interactive path vector stacking. The deep learning network is used to predict the fault root cause description label to achieve accurate identification and positioning of fault activities.
It improves the targetedness and positioning speed of fault identification, enhances the depth and accuracy of fault analysis, and provides efficient fault management support.
Smart Images

Figure CN119629029B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of Internet of Things operation and maintenance, and in particular, to a distributed operation and maintenance intelligent management method and system based on cloud computing and the Internet of Things. Background Art
[0002] With the rapid development of cloud computing and Internet of Things technologies, distributed systems have become the mainstream architecture for information processing and service provision. In these complex distributed environments, the interaction event streams between Internet of Things devices are becoming increasingly large and complex, resulting in unprecedented challenges for operation and maintenance management. Traditional operation and maintenance methods often rely on manual experience and simple monitoring tools, and it is difficult to effectively address the problems of fault location and management in large-scale and highly dynamic Internet of Things interaction event streams.
[0003] Especially in a cloud computing environment, the interaction event streams between Internet of Things devices involve numerous potential fault activities, and these fault activities are often interrelated and interact with each other, forming complex fault chains. In order to accurately locate and solve these faults, operation and maintenance personnel need to deeply understand the logical relationships and propagation paths of fault activities, which is often time-consuming, laborious and error-prone in actual operations. Summary of the Invention
[0004] In view of the problems mentioned above, in combination with the first aspect of the present invention, embodiments of the present invention provide a distributed operation and maintenance intelligent management method based on cloud computing and the Internet of Things, and the method includes:
[0005] Obtain the inference fault knowledge chains corresponding to each potential fault activity in the target Internet of Things interaction event stream in the cloud computing environment, and determine the concerned fault knowledge chain among the multiple inference fault knowledge chains, where the concerned fault knowledge chain is used to guide the location of the concerned target fault activity among the multiple potential fault activities;
[0006] Perform encoding representations of the target Internet of Things interaction event stream respectively according to the concerned fault knowledge chain and each of the inference fault knowledge chains, generate multiple fault logic graph encoding vectors, and extract the fault penetration node vectors corresponding to the concerned fault knowledge chain and each of the inference fault knowledge chains;
[0007] Stack each of the fault logic graph encoding vectors with the corresponding fault penetration node vector to generate multiple fault logic transfer vectors;
[0008] Extract the first interaction path vector of the target Internet of Things interaction event stream, and stack the first interaction path vector and the multiple fault logic transfer vectors to generate a target stack vector;
[0009] Perform prediction according to the target stack vector to generate a target fault root cause description label of the target fault activity.
[0010] In a possible implementation manner of the first aspect, the stacking of the first interaction path vector and the multiple fault logic transfer vectors to generate a target stacking vector includes:
[0011] Determine the knowledge chain lengths of the respective inference fault knowledge chains, and stack the fault logic transfer vectors corresponding to the respective inference fault knowledge chains in descending order of the knowledge chain lengths to generate a first stacking vector;
[0012] Stack the first interaction path vector, the first stacking vector, and the fault logic transfer vector corresponding to the concerned fault knowledge chain to generate a target stacking vector.
[0013] In a possible implementation manner of the first aspect, the target fault root cause description label is generated by a first deep learning network. The stacking of the first interaction path vector, the first stacking vector, and the fault logic transfer vector corresponding to the concerned fault knowledge chain to generate a target stacking vector includes:
[0014] Generate guiding information for guiding the first deep learning network to generate a fault root cause description label;
[0015] Extract the guiding knowledge vector of the guiding information, and stack the first interaction path vector, the guiding knowledge vector, the first stacking vector, and the fault logic transfer vector corresponding to the concerned fault knowledge chain to generate a target stacking vector.
[0016] In a possible implementation manner of the first aspect, the encoding representations of the target Internet of Things interaction event stream according to the concerned fault knowledge chain and each of the inference fault knowledge chains respectively to generate multiple fault logic diagram encoding vectors include:
[0017] Perform multi-stage encoding representation on the target Internet of Things interaction event stream to generate multi-stage encoding features of the target Internet of Things interaction event stream;
[0018] According to the concerned fault knowledge chain and each of the inference fault knowledge chains respectively, perform knowledge chain encoding mapping on the multi-stage encoding features to generate multiple multi-stage mapping encoding features;
[0019] Perform encoding integration on each of the multi-stage mapping encoding features respectively to generate multiple fault logic diagram encoding vectors.
[0020] In a possible implementation of the first aspect, the obtaining of the inference failure knowledge chains corresponding to each potential failure activity in the target Internet of Things interaction event stream in the cloud computing environment and the determination of the concerned failure knowledge chain among the multiple inference failure knowledge chains include:
[0021] Obtain the target Internet of Things interaction event stream and the guiding identification information of the target Internet of Things interaction event stream. The target Internet of Things interaction event stream includes multiple potential failure activities, and the guiding identification information is used to guide the positioning of the target failure activity concerned among the multiple potential failure activities;
[0022] Perform failure knowledge chain extraction on the target Internet of Things interaction event stream to generate the inference failure knowledge chains corresponding to each potential failure activity, and determine the concerned failure knowledge chain among the inference failure knowledge chains according to the guiding identification information.
[0023] In a possible implementation of the first aspect, the step of performing failure knowledge chain extraction on the target Internet of Things interaction event stream to generate the inference failure knowledge chains corresponding to each potential failure activity includes:
[0024] Traverse the target Internet of Things interaction event stream, add each event in the target Internet of Things interaction event stream as a node to the event association graph, and determine the association relationship between events according to the time sequence, event source and event target attributes between events. After adding the corresponding edges to the event association graph, optimize the event association graph;
[0025] Use the graph traversal algorithm to traverse the nodes in the optimized event association graph, identify potential failure activities according to the association relationship between events and the preset failure identification rules, represent the identified failure activities in the form of a node set, and add them to the list of potential failure activities;
[0026] For each potential failure activity in the list of potential failure activities, extract the corresponding key features, and the key features include event type distribution, event time interval distribution, event source and target distribution, etc.;
[0027] Quantify the key features into numerical data, and after combining them into a failure feature vector, perform normalization processing on the failure feature vector;
[0028] For the failure feature vector of each potential failure activity, use the similarity calculation algorithm to find the most similar failure feature in the pre-constructed failure knowledge base, obtain the corresponding failure knowledge data according to the found failure feature, and construct an inference failure knowledge chain according to the association relationship between the failure types and the logical relationship between the failure causes in the failure knowledge data;
[0029] For example, in a possible implementation of the first aspect, the steps of respectively performing encoding representations of the target Internet of Things interaction event stream according to the concerned fault knowledge chain and each of the inferred fault knowledge chains, generating a plurality of fault logic diagram encoding vectors, and extracting the fault penetration node vectors corresponding to the concerned fault knowledge chain and each of the inferred fault knowledge chains include:
[0030] Uniquely identify each event in the target Internet of Things interaction event stream, and construct an initial empty matrix as the event stream - knowledge chain association matrix;
[0031] Traverse each event in the target Internet of Things interaction event stream. For each event, analyze whether the event appears in the concerned fault knowledge chain or any of the inferred fault knowledge chains. If so, mark the corresponding position in the event stream - knowledge chain association matrix as 1, otherwise mark it as 0;
[0032] Adopt a time window division or event type clustering method to divide the target Internet of Things interaction event stream into multiple stages. For each stage, extract the key features of the events in this stage, and encode the key features into numerical data. After forming the multi - stage encoding features of this stage, combine the multi - stage encoding features of all stages in sequence to obtain the overall multi - stage encoding features of the target Internet of Things interaction event stream. The key features include event type, event frequency, and time interval between events;
[0033] For each inferred fault knowledge chain, determine the knowledge chain nodes associated with the events in the target Internet of Things interaction event stream according to the event stream - knowledge chain association matrix;
[0034] For each associated knowledge chain node, perform knowledge chain encoding mapping according to the position and role of the knowledge chain node in the knowledge chain and the multi - stage encoding features of the corresponding event. The mapping process uses the method of neural network mapping to transform the multi - stage encoding features of the corresponding event into mapping encoding features corresponding to the knowledge chain node. And for each inferred fault knowledge chain, combine the mapping encoding features of all associated knowledge chain nodes in sequence to form the multi - stage mapping encoding features corresponding to this inferred fault knowledge chain;
[0035] Input each multi - stage mapping encoding feature into a recurrent neural network so that the recurrent neural network generates fault logic diagram encoding vectors by learning and extracting the association relationships and temporal information between each multi - stage mapping encoding feature;
[0036] Analyze the fault propagation path and fault penetration nodes of each inferred fault knowledge chain. The fault penetration node refers to the key node that can affect or trigger other faults when the fault propagates in the inferred fault knowledge chain;
[0037] According to the position, role of the fault penetration node in the inference fault knowledge chain and the concerned fault knowledge chain, and its association relationship with other fault penetration nodes, extract the key features of the fault penetration node, and encode the key features as numerical data to form a fault penetration node vector. The key features include node type, node degree, and path length between nodes.
[0038] For each inference fault knowledge chain, stack the corresponding fault logic graph encoding vector and the fault penetration node vector to obtain the fault logic transfer vector corresponding to the inference fault knowledge chain. The attention mechanism or the gating mechanism is used in the stacking process to dynamically adjust the weights of the fault logic graph encoding vector and the fault penetration node vector.
[0039] In a possible implementation manner of the first aspect, the target fault root cause description label is generated by a first deep learning network. The training steps of the first deep learning network include:
[0040] Obtain the template Internet of Things interaction event stream and the second fault knowledge chain corresponding to the template fault activity in the template Internet of Things interaction event stream. Extract the fault knowledge chain from the template Internet of Things interaction event stream to generate the first fault knowledge chain corresponding to each fault interaction activity in the template Internet of Things interaction event stream. The template fault activity is one of the multiple fault interaction activities.
[0041] Extract the second interaction path vector of the template Internet of Things interaction event stream, perform encoding representations of the template Internet of Things interaction event stream respectively according to the second fault knowledge chain and each of the first fault knowledge chains to generate multiple template encoding features, and extract the template fault penetration node vectors corresponding to the second fault knowledge chain and each of the first fault knowledge chains.
[0042] Stack the second interaction path vector, the template encoding features, and the template fault penetration node vectors and load them into the first deep learning network for prediction to generate a fault root cause description prediction result, which is used to determine the fault root cause description label of the template fault activity.
[0043] Obtain the first fault root cause description label of the template fault knowledge point associated with the template Internet of Things interaction event stream. Determine the training error according to the fault root cause description prediction result and the first fault root cause description label, and optimize the first deep learning network according to the training error. The template fault knowledge point is used to guide the positioning of the template fault activity.
[0044] In a possible implementation of the first aspect, the template Internet of Things interaction event stream, the second fault knowledge chain, and the first fault root cause description label are all obtained from a sample database. Before obtaining the template Internet of Things interaction event stream and the second fault knowledge chain corresponding to the template fault activity in the template Internet of Things interaction event stream, the method further includes:
[0045] Obtain a plurality of basic Internet of Things interaction event streams and the fault attention knowledge data corresponding to each of the basic Internet of Things interaction event streams. Based on each of the basic Internet of Things interaction event streams and the corresponding fault attention knowledge data, respectively determine the key interaction point data corresponding to each of the basic Internet of Things interaction event streams. The fault attention knowledge data is used to guide the positioning of the key fault activities in the corresponding basic Internet of Things interaction event stream;
[0046] Obtain a plurality of reference fault knowledge points. According to each of the key interaction point data, respectively determine the associated fault knowledge points corresponding to each of the basic Internet of Things interaction event streams among the plurality of reference fault knowledge points;
[0047] According to each of the basic Internet of Things interaction event streams and the corresponding fault attention knowledge data, respectively determine the fault knowledge chain label data corresponding to each of the basic Internet of Things interaction event streams. The fault knowledge chain label data is used to guide the positioning of the corresponding key fault activities;
[0048] Integrate and add each of the basic Internet of Things interaction event streams, the corresponding fault knowledge chain label data, and the corresponding associated fault knowledge points to the sample database. The template Internet of Things interaction event stream is randomly selected from the plurality of basic Internet of Things interaction event streams. The second fault knowledge chain is the fault knowledge chain label data corresponding to the template Internet of Things interaction event stream, and the template fault knowledge point is the associated fault knowledge point associated with the template Internet of Things interaction event stream.
[0049] In a possible implementation of the first aspect, the step of respectively determining the fault knowledge chain label data corresponding to each of the basic Internet of Things interaction event streams according to each of the basic Internet of Things interaction event streams and the corresponding fault attention knowledge data includes:
[0050] Load each of the fault attention knowledge data into a second deep learning network for prediction to generate the fault summary description data corresponding to each of the basic Internet of Things interaction event streams;
[0051] Fault activity localization is performed respectively based on each of the basic Internet of Things interaction event streams and the corresponding fault summary description data, to generate basic fault logic paths corresponding to each of the basic Internet of Things interaction event streams, where the basic fault logic paths are used to guide the localization of the corresponding key fault activities;
[0052] Each of the basic Internet of Things interaction event streams and the corresponding basic fault logic paths are respectively loaded into a first fault knowledge chain extraction network for fault knowledge chain extraction, to generate fault knowledge chain label data corresponding to each of the basic Internet of Things interaction event streams.
[0053] On the other hand, an embodiment of the present invention further provides a distributed operation and maintenance intelligent management system based on cloud computing and the Internet of Things, including a processor and a machine-readable storage medium, where the machine-readable storage medium is connected to the processor, the machine-readable storage medium is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the machine-readable storage medium to implement the above method.
[0054] Based on the above aspects, the embodiments of the present application achieve efficient fault management and localization of Internet of Things interaction event streams in a cloud computing environment. First, the inference fault knowledge chain of potential fault activities is accurately obtained and screened, the target fault activities to be concerned are determined, and the pertinence of fault identification is improved. By encoding and representing the target Internet of Things interaction event stream, generating a fault logic graph encoding vector and a fault penetration node vector, and further stacking to form a fault logic transfer vector, the fault knowledge and event stream features are effectively fused, enhancing the depth and accuracy of fault analysis. Finally, by combining the first interaction path vector of the target Internet of Things interaction event stream with the fault logic transfer vector, a target stacking vector is generated, and based on this, the fault root cause description label of the target fault activity is predicted, significantly improving the speed and accuracy of fault localization, and providing data support for operation and maintenance management in cloud computing and Internet of Things environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 is a schematic execution flow diagram of a distributed operation and maintenance intelligent management method based on cloud computing and the Internet of Things provided by an embodiment of the present invention.
[0056] Figure 2 is a schematic hardware architecture diagram of a distributed operation and maintenance intelligent management system based on cloud computing and the Internet of Things provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0057] The present invention will be specifically described below with reference to the accompanying drawings of the specification, Figure 1It is a schematic flowchart of a distributed operation and maintenance intelligent management method based on cloud computing and the Internet of Things provided by an embodiment of the present invention. The distributed operation and maintenance intelligent management method based on cloud computing and the Internet of Things will be introduced in detail below.
[0058] Step S110, obtain the inference failure knowledge chains corresponding to each potential failure activity in the target Internet of Things interaction event stream in the cloud computing environment, and determine the concerned failure knowledge chain among the multiple inference failure knowledge chains. The concerned failure knowledge chain is used to guide the positioning of the target failure activity concerned among the multiple potential failure activities.
[0059] In this embodiment, taking the intelligent operation and maintenance scenario of radio monitoring facilities based on distributed mobile terminals as an example, radio monitoring facilities include numerous Internet of Things devices, and these Internet of Things devices interact with each other to achieve monitoring functions. For example, sensor devices transmit the collected radio signal data to the data processing unit, and the data processing unit then interacts with the storage device for data storage, etc. The cloud computing environment bears the event stream generated by the interaction of these Internet of Things devices.
[0060] The server obtains the target Internet of Things interaction event stream in the cloud computing environment. In this target Internet of Things interaction event stream, the acquisition process of the inference failure knowledge chains corresponding to multiple potential failure activities is included. Suppose there is a sensor node in the radio monitoring facility, and under normal circumstances, it will send the monitored radio frequency band data to the data aggregation node at regular time intervals. If the data transmission from this sensor node is not received for a period of time, this may be a potential failure activity.
[0061] For this potential failure activity, the server constructs an inference failure knowledge chain by analyzing historical data, device operation manuals, known failure modes, etc. This knowledge chain may include a series of possible failure factors and their associations from device hardware failures (such as sensor element damage) to software failures (such as data transmission protocol errors), and then to network-level failures (such as network connection interruption between the sensor and the aggregation node).
[0062] Similarly, for other potential failure activities, such as the data processing unit processing data too slowly, the storage device failing to write data, etc., the server also constructs corresponding inference failure knowledge chains.
[0063] Then, determine the concerned fault knowledge chain among numerous inference fault knowledge chains. For example, if the current operation and maintenance personnel feedback that the accuracy of monitoring data in a specific area has severely declined, the server, based on this guiding identification information (corresponding to the problem of the decline in the accuracy of monitoring data in this specific area), determines the concerned fault knowledge chain related to this problem among each inference fault knowledge chain. If it is the decline in the accuracy of monitoring data, the concerned fault knowledge chain may focus on the fault factors and their correlation relationships in aspects such as the accuracy of sensor-collected data, data integrity during data transmission, and data calibration in the data processing unit. This concerned fault knowledge chain will guide the positioning of the target fault activities among multiple potential fault activities, that is, those fault activities directly or indirectly related to the decline in the accuracy of monitoring data.
[0064] Step S120, perform encoding representations on the target Internet of Things interaction event stream respectively according to the concerned fault knowledge chain and each of the inference fault knowledge chains, generate multiple fault logic diagram encoding vectors, and extract the fault penetration node vectors corresponding to the concerned fault knowledge chain and each of the inference fault knowledge chains.
[0065] In this embodiment, for each event in the target Internet of Things interaction event stream, such as an event where a sensor collects a radio signal strength of a certain value, or an event where the data processing unit takes a specific duration to process a batch of data, etc., encoding representation needs to be performed.
[0066] The server first performs multi-stage encoding representation on the target Internet of Things interaction event stream. For example, according to the time sequence of event occurrence, the entire interaction event stream is divided into multiple time stages, and the key features of the events are encoded within each stage. For sensor collection events, the key features may include collection time, signal strength, signal frequency, etc.; for data processing events, the key features may include the amount of processed data, processing duration, processing algorithm, etc. After these key features are encoded into numerical data, the multi-stage encoding features of the target Internet of Things interaction event stream are generated.
[0067] Then, respectively according to the concerned fault knowledge chain and each inference fault knowledge chain, perform knowledge chain encoding mapping on the multi-stage encoding features. Taking the concerned fault knowledge chain as an example, if the concerned fault knowledge chain emphasizes the accuracy of sensor data, then for the multi-stage encoding features of sensor collection events, mapping will be performed according to the fault knowledge chain rules related to the accuracy of sensor data. For example, if the knowledge chain indicates that abnormal fluctuations in signal strength are related to sensor aging, then in the encoding mapping process, the feature of signal strength will have a specific mapping encoding in the dimension related to sensor aging.
[0068] Then, encode and integrate each multi-stage mapping encoding feature respectively to generate multiple fault logic diagram encoding vectors. For example, integrate the multi-stage mapping encoding features related to each link such as sensor acquisition, data processing, and data storage according to certain logic and algorithms to obtain an encoding vector that can comprehensively reflect the fault logic relationship.
[0069] Meanwhile, the server also needs to extract the concerned fault knowledge chain and the fault penetration node vectors corresponding to each inferred fault knowledge chain. In a radio monitoring facility, the fault penetration nodes may be those nodes that will affect multiple subsequent links once a fault occurs. For example, the data verification module in the data processing unit, if it malfunctions, may cause incorrect data to flow into the storage device, thus affecting the accuracy of the entire monitoring data. For this data verification module, its key features include the module type (data verification), the connection relationship with other modules (receiving the data processed by the sensor and outputting the data to the storage device), etc. Encoding these key features into numerical data forms the fault penetration node vector.
[0070] Step S130, stack each of the fault logic diagram encoding vectors with the corresponding fault penetration node vector respectively to generate multiple fault logic transfer vectors.
[0071] Taking the sensor acquisition link as an example, assume that the previously generated fault logic diagram encoding vector related to sensor acquisition contains the fault logic information of all aspects of sensor acquisition, and the corresponding fault penetration node vector contains the key information of the fault penetration node such as the sensor element.
[0072] The server stacks these two vectors through a specific algorithm (such as weighted summation, etc.). This stacking process is like integrating the fault logic information in the sensor acquisition link with the key node information that may cause fault propagation to form a fault logic transfer vector that more comprehensively reflects the fault transfer situation in the sensor acquisition link. Similarly, for other links such as data processing and data storage, stack their respective fault logic diagram encoding vectors with the fault penetration node vectors to generate the fault logic transfer vectors of their respective links.
[0073] Step S140, extract the first interaction path vector of the target Internet of Things interaction event stream, and stack the first interaction path vector and the multiple fault logic transfer vectors to generate a target stacked vector.
[0074] In this embodiment, this first interaction path vector reflects the sequence and path relationship of the interaction between Internet of Things devices. For example, the sensor transmits data to the data processing unit, and the data processing unit then transmits the processed data to the storage device. This transmission sequence and the device relationship involved constitute the first interaction path vector.
[0075] Then, the server stacks the first interaction path vector and multiple fault logic flow vectors to generate a target stack vector. For example, first stack the fault logic flow vectors corresponding to each inference fault knowledge chain in descending order of the knowledge chain length of each inference fault knowledge chain to generate a first stack vector. If among the previous fault logic flow vectors, the fault logic flow vector related to the data processing unit has a longer knowledge chain length, it may be stacked first.
[0076] Next, stack the first interaction path vector, the first stack vector, and the fault logic flow vector corresponding to the concerned fault knowledge chain. This process is like integrating the overall path information of device interaction, the fault flow information of each link, and the fault flow information of the key - concerned fault link together to form a target stack vector.
[0077] Step S150, make a prediction based on the target stack vector to generate a target fault root cause description label for the target fault activity.
[0078] For example, if the target fault activity is the decline in the accuracy of monitored data, the target fault root cause description label may include a detailed fault analysis starting from the sensor acquisition link. It may describe the external interference suffered by the sensor when collecting data (such as signal interference from other nearby radio devices), how this interference causes deviation in the data collected by the sensor, and when the deviated data passes through the data processing unit, due to a certain algorithm in the data processing unit not effectively correcting this deviation (possibly due to the limitations of the algorithm itself or unreasonable algorithm parameter settings), the processed data still has a large error, and then these incorrect data are stored in the storage device, ultimately leading to the decline in the accuracy of monitored data. At the same time, this target fault root cause description label may also include the impact of data packet loss during network transmission on data accuracy, and the impact of time synchronization problems in the entire Internet of Things interaction system on data accuracy, etc., including complex logical relationships and detailed fault factor descriptions in multiple aspects.
[0079] Thus, in the intelligent operation and maintenance scenario of radio monitoring facilities based on distributed mobile terminals, it is possible to accurately analyze the fault activities in the target Internet of Things interaction event stream, find the root cause of the target fault activity, and generate a detailed target fault root cause description label.
[0080] Based on the above steps, the embodiments of the present application achieve efficient fault management and localization for the Internet of Things interaction event stream in the cloud computing environment. First, it accurately obtains and filters the inference fault knowledge chain of potential fault activities, determines the target fault activities of concern, and improves the pertinence of fault identification. By encoding and representing the target Internet of Things interaction event stream, generating the fault logic graph encoding vector and the fault penetration node vector, and further stacking them to form the fault logic transfer vector, it effectively integrates the fault knowledge and the event stream characteristics, enhancing the depth and accuracy of fault analysis. Finally, by combining the first interaction path vector of the target Internet of Things interaction event stream with the fault logic transfer vector, generating the target stacking vector, and predicting the fault root cause description label of the target fault activity based on this, it significantly improves the speed and accuracy of fault localization, providing data support for the operation and maintenance management in the cloud computing and Internet of Things environments.
[0081] In a possible implementation manner, step S140 includes:
[0082] Step S141, respectively determine the knowledge chain lengths of each of the inference fault knowledge chains, and stack the fault logic transfer vectors corresponding to each of the inference fault knowledge chains in descending order of the knowledge chain lengths to generate a first stacking vector.
[0083] Step S142, stack the first interaction path vector, the first stacking vector, and the fault logic transfer vector corresponding to the concerned fault knowledge chain to generate a target stacking vector.
[0084] In a possible implementation manner, the target fault root cause description label is generated by a first deep learning network, and step S142 includes:
[0085] Step S1421, generate guiding information for guiding the first deep learning network to generate the fault root cause description label.
[0086] Step S1422, extract the guiding knowledge vector of the guiding information, and stack the first interaction path vector, the guiding knowledge vector, the first stacking vector, and the fault logic transfer vector corresponding to the concerned fault knowledge chain to generate a target stacking vector.
[0087] In this embodiment, the length of each inference fault knowledge chain is first determined. For example, taking the interaction between the aforementioned sensor, data processing unit, and storage device as an example, different fault inference knowledge chains may exist in each link. For the sensor link, its inference fault knowledge chain may start from the sensor hardware itself (such as component aging, damage, etc.), extend to its connection with peripheral devices (such as loose connection lines, interface faults, etc.), and then to various aspects such as possible external interference (such as electromagnetic interference, etc.), which constitutes a relatively long knowledge chain. For the data processing unit, its inference fault knowledge chain may be more concentrated in internal algorithms (such as algorithm errors, improper parameter settings, etc.), software operating environment (such as operating system faults affecting data processing), etc., and relatively speaking, the knowledge chain length may be shorter.
[0088] Assume that the inference fault knowledge chain of the sensor link is the longest, then the fault logic transfer vector related to the sensor will be stacked first. This fault logic transfer vector contains all the logical information from data acquisition in the sensor link to possible faults and their propagation. When stacking in descending order of the knowledge chain length, the fault logic transfer vectors of subsequent links such as the data processing unit and storage device are stacked in turn. This process is like integrating the fault logic transfer information of each link from the most complex and comprehensive fault logical relationship to the relatively simple one to form the first stacked vector.
[0089] The first interaction path vector reflects the interaction sequence between IoT devices in the radio monitoring facility. For example, the complete interaction process information that after the sensor collects data and transmits it to the data processing unit, and then the data processing unit transmits the processed data to the storage device. Stack this interaction path vector with the previously generated first stacked vector and the fault logic transfer vector corresponding to the concerned fault knowledge chain. If the concerned fault knowledge chain is centered around data accuracy, then its corresponding fault logic transfer vector contains the fault logical relationships that may affect data accuracy in each link from sensor acquisition to final storage related to data accuracy. Through such stacking, the target stacked vector integrates information such as the device interaction path, the fault logic transfer of each link, and the key concerned fault logical relationships.
[0090] When the target fault root cause description label is generated by the first deep learning network, during the process of generating the target stacked vector, the server first needs to generate guiding information for guiding the first deep learning network to generate the fault root cause description label. In the intelligent operation and maintenance scenario of radio monitoring facilities, the guiding information may come from the preliminary description of the fault phenomenon by the operation and maintenance personnel, specific monitoring index requirements, or key performance parameters of equipment operation, etc. For example, the operation and maintenance personnel feedback that the radio signal data in a specific frequency band fluctuates abnormally during a certain period, and this information is part of the guiding information. At the same time, the monitoring facilities have specific index requirements for data accuracy, data integrity, etc., and these requirements also constitute the guiding information.
[0091] Among them, the server then extracts the guiding knowledge vector of the guiding information. This guiding knowledge vector is obtained by quantifying and encoding various contents in the guiding information. For example, for the information of abnormal signal data fluctuation feedback by the operation and maintenance personnel, the amplitude range of the fluctuation, the frequency of the fluctuation, etc. can be quantified and encoded. For the index requirements of the monitoring facilities, such as the data accuracy requirement reaching more than 99%, this 99% and other related index values can be encoded. Then, the first interaction path vector, the guiding knowledge vector, the first stacked vector, and the fault logic transfer vector corresponding to the concerned fault knowledge chain are stacked to generate the target stacked vector. In this process, the addition of the guiding knowledge vector is like injecting specific target-oriented information into the entire stacking process. The first interaction path vector provides the basic framework of device interaction, the first stacked vector contains the fault logic transfer situation of each link, the fault logic transfer vector corresponding to the concerned fault knowledge chain focuses on the key concerned fault logic relationships, and the guiding knowledge vector guides the deep learning network to generate the fault root cause description label in the direction that meets the operation and maintenance requirements and the characteristics of the fault phenomenon. Through such stacking, the target stacked vector integrates all necessary information so that the subsequent first deep learning network can accurately generate a complex target fault root cause description label based on this target stacked vector. This label will detail the root cause of the fault and the logical relationship between various factors during the process from sensor data collection to data storage.
[0092] In a possible implementation manner, step S120 may include:
[0093] Perform multi-stage coding representation on the target Internet of Things interaction event stream to generate multi-stage coding features of the target Internet of Things interaction event stream.
[0094] Respectively, according to the concerned fault knowledge chain and each of the inferred fault knowledge chains, perform knowledge chain coding mapping on the multi-stage coding features to generate multiple multi-stage mapping coding features.
[0095] Encode and integrate each of the multi-stage mapping coding features to generate multiple fault logic diagram coding vectors.
[0096] In this embodiment, the Internet of Things interaction event stream covers the interaction processes among numerous devices, such as the interaction between sensors and data processing units, data processing units and storage devices, etc. The server divides this interaction event stream into multiple stages according to the time sequence or a specific task process. For example, taking the case where a sensor collects radio signals and transmits them to a data processing unit as an example, the initial calibration stage after the sensor is started can be set as the first stage, the stage of normal data collection and transmission as the second stage, and the adjustment stage when external interference is encountered as the third stage, etc.
[0097] For the events in each stage, the server extracts their key features for encoding. In the initial calibration stage of the sensor, the key features may include the startup time of the sensor, initial self-check parameters (such as voltage, current, etc.), the version of the calibration algorithm, etc. After these features are quantified into numerical data, they constitute the coding representation of this stage. In the normal data collection and transmission stage, the key features are the intensity, frequency, data collection time interval, data transmission rate, etc. of the collected radio signals. For the adjustment stage when external interference is encountered, the key features may be the time when the interference is detected, the type of interference (such as the frequency band of electromagnetic interference, etc.), the adjustment strategy adopted by the sensor (such as the parameters corresponding to operations such as adjusting the gain), etc. Combining the encodings of these stages generates the multi-stage coding features of the target Internet of Things interaction event stream.
[0098] Next, the server respectively performs knowledge chain coding mapping on the multi-stage coding features according to the concerned fault knowledge chain and each inferred fault knowledge chain to generate multiple multi-stage mapping coding features. In the intelligent operation and maintenance scenario of radio monitoring facilities, if the concerned fault knowledge chain focuses on data accuracy, each element in the multi-stage coding features should be mapped according to this concerned fault knowledge chain. Taking the data collection stage of the sensor as an example, for the element of the intensity of the collected radio signal in the multi-stage coding features, if the concerned fault knowledge chain indicates that abnormal signal intensity may be related to sensor aging, then in the knowledge chain coding mapping process, this signal intensity value will be mapped to the coding dimension related to sensor aging. If the signal intensity is within the normal range, it may be mapped to a value representing the normal aging state. If the signal intensity is too low and exceeds the normal fluctuation range, it may be mapped to a value representing that sensor aging seriously affects the data collection accuracy.
[0099] The same applies to each inference failure knowledge chain. For example, there is an inference failure knowledge chain regarding the impact of network transmission on data accuracy. When encoding and mapping the knowledge chain, for the element of data transmission rate in the multi-stage encoded features, if the transmission rate is lower than the normal level, according to this inference failure knowledge chain, it may be mapped to the encoding dimension representing network congestion or network device failures affecting data transmission and thus data accuracy. In this way, each element in each multi-stage encoded feature is mapped based on the concerned failure knowledge chain and each inference failure knowledge chain, thereby generating multiple multi-stage mapped encoded features.
[0100] Finally, the server encodes and integrates each multi-stage mapped encoded feature respectively to generate multiple fault logic diagram encoding vectors. In a radio monitoring facility, taking the operation and maintenance of sensors, data processing units, and storage devices as an example, for the multi-stage mapped encoded features related to sensors, the server integrates the encoded results after mapping at different stages according to specific logics and algorithms. For example, weighted summation or combination according to specific logic rules is performed on the multi-stage mapped encoded features related to sensors, such as the initial calibration stage, normal acquisition and transmission stage, and adjustment stage in case of external interference. During this process, different weights may be assigned according to the importance of different stages to the overall fault logic. For example, the weight of the normal acquisition and transmission stage on the overall fault logic may be relatively high because this stage has a long duration and the accuracy of data acquisition is directly related to the operation effect of the entire monitoring facility.
[0101] Similarly, similar encoding and integration operations are performed on the multi-stage mapped encoded features related to data processing units and storage devices. Through such an integration process, the server generates multiple fault logic diagram encoding vectors, which can comprehensively reflect the fault logic relationships of each link (such as sensors, data processing units, storage devices, etc.) based on the concerned failure knowledge chain and inference failure knowledge chain in the entire Internet of Things interaction event stream, providing an important data basis for subsequent fault analysis and root cause location.
[0102] In a possible implementation manner, step S110 includes:
[0103] Step S111, obtaining a target Internet of Things interaction event stream and the guiding identification information of the target Internet of Things interaction event stream, where the target Internet of Things interaction event stream includes multiple potential fault activities, and the guiding identification information is used to guide the positioning of the target fault activity of concern among the multiple potential fault activities.
[0104] Step S112: Extract the fault knowledge chain from the target Internet of Things interaction event stream to generate the inference fault knowledge chain corresponding to each potential fault activity, and determine the concerned fault knowledge chain in each inference fault knowledge chain according to the guiding identification information.
[0105] In a possible implementation manner, step S112 includes:
[0106] Step S1121: Traverse the target Internet of Things interaction event stream, add each event in the target Internet of Things interaction event stream as a node to the event association graph, and determine the association relationship between events according to the time sequence, event source, and event target attributes between events. After adding the corresponding edges to the event association graph, optimize the event association graph.
[0107] Step S1122: Use the graph traversal algorithm to traverse the nodes in the optimized event association graph, identify potential fault activities according to the association relationship between events and the preset fault recognition rules, represent the identified fault activities in the form of a node set, and add them to the potential fault activity list.
[0108] Step S1123: For each potential fault activity in the potential fault activity list, extract the corresponding key features, where the key features include event type distribution, event time interval distribution, event source and target distribution, etc.
[0109] Step S1124: Quantify the key features into numerical data, combine them into a fault feature vector, and then perform normalization processing on the fault feature vector.
[0110] Step S1125: For the fault feature vector of each potential fault activity, use the similarity calculation algorithm to find the most similar fault feature in the pre-constructed fault knowledge base, obtain the corresponding fault knowledge data according to the found fault feature, and construct an inference fault knowledge chain according to the association relationship between fault types and the logical relationship between fault causes in the fault knowledge data.
[0111] In this embodiment, the target Internet of Things interaction event stream contains numerous interaction events between devices. For example, sensor devices send radio signal data to the data processing unit, and the data processing unit processes the data and stores it in the storage device, etc. The guiding identification information may come from the requirements of operation and maintenance personnel or specific monitoring indicators of the monitoring system. For example, operation and maintenance personnel find that the accuracy of recent radio monitoring data has decreased. This information about the decrease in data accuracy is the guiding identification information, which will guide the server to locate the target fault activity related to the decrease in data accuracy among numerous potential fault activities.
[0112] Next, during the process of extracting the fault knowledge chain, the target Internet of Things interaction event stream can be traversed, and each event therein is added as a node to the event association graph. Taking the event that the sensor sends data to the data processing unit as an example, the event that the sensor sends data is a node, and the event that the data processing unit receives data is also a node. Then, the association relationship between events is determined according to the chronological order, event source, and event target attributes between events, and corresponding edges are added to the event association graph. If the sensor sends data first and the data processing unit receives data later, then an edge is established from the node where the sensor sends data to the node where the data processing unit receives data according to this chronological order. Moreover, since the sensor is the event source and the data processing unit is the event target, this also clarifies the direction of the edge. After that, the server will optimize this event association graph, and may perform operations such as removing some redundant edges or merging some nodes with similar attributes.
[0113] After the optimized event association graph is formed, the server traverses the nodes therein using a graph traversal algorithm. Potential fault activities are identified according to the association relationship between events and the preset fault recognition rules. For example, according to the preset rules, if the data sent by the sensor continuously exceeds the normal data range for multiple times, this may be a potential fault activity. All identified fault activities are represented in the form of a node set and added to the list of potential fault activities.
[0114] For each potential fault activity in the list of potential fault activities, the server extracts the corresponding key features. In the scenario of radio monitoring facilities, taking the potential fault activity of abnormal sensor data transmission as an example, in terms of the event type distribution, it may be that the data transmission type is incorrect (for example, data in a specific frequency band should be sent but data in other frequency bands is sent); in terms of the event time interval distribution, it may be that the time interval for sending data does not match the normal setting (for example, data is normally sent every 1 second, but now it is sent every 5 seconds); in terms of the event source and target distribution, the source is the sensor and the target is the data processing unit. If the data always fails to reach the target accurately, this is also one of the key features.
[0115] Then, the server quantifies these key features into numerical data, combines them into a fault feature vector, and then normalizes the fault feature vector. For example, the incorrect data transmission type is converted into a numerical value according to the preset coding rule. For example, the type error is encoded as 1 and the normal is 0; for the time interval, the difference from the normal time interval is calculated and quantified into a numerical value; after these numerical values are combined into a fault feature vector, the numerical values in the vector are mapped to a specific interval, such as the interval [0,1], through a normalization algorithm.
[0116] Finally, for the fault feature vector of each potential fault activity, the server uses a similarity calculation algorithm to find the most similar fault feature in the pre-constructed fault knowledge base. In the fault knowledge base of the radio monitoring facility, the feature data of various known faults are stored. For example, for the fault feature vector of abnormal sensor data transmission, the server will search for a similar fault feature in the knowledge base. If a similar fault feature is found, the corresponding fault knowledge data may include the fault type (such as sensor hardware failure or software configuration error) when this fault occurred before and the logical relationship between the fault causes (such as hardware failure causing a change in data transmission frequency or software configuration error causing an error in data transmission type). Based on the association relationship between these fault types and the logical relationship between the fault causes, the server constructs an inferential fault knowledge chain. This inferential fault knowledge chain may start from a component failure of the sensor hardware and deduce how it affects the data transmission type or time interval of the sensor, and then affect a series of logical relationships such as the accuracy of the data received by the data processing unit.
[0117] After constructing the inferential fault knowledge chains corresponding to each potential fault activity, the server determines the concerned fault knowledge chain among these inferential fault knowledge chains based on the previously obtained guiding identification information. Since the guiding identification information is about the decrease in data accuracy, among the various inferential fault knowledge chains, the inferential fault knowledge chains related to data accuracy will be determined as the concerned fault knowledge chains. For example, those inferential fault knowledge chains that contain logical relationships such as inaccurate sensor-acquired data affecting the accuracy of subsequent data processing and algorithm errors in the data processing unit leading to a decrease in data accuracy will be focused on. This concerned fault knowledge chain will play an important role in guiding and positioning the target fault activity in the subsequent fault analysis process.
[0118] For example, in a possible implementation manner, step S120 may further include:
[0119] Step S121, uniquely identify each event in the target Internet of Things interaction event stream and construct an initial empty matrix as the event stream - knowledge chain association matrix.
[0120] In this embodiment, the Internet of Things interaction event stream contains numerous events, such as events of the sensor collecting radio signals, events of the data processing unit receiving sensor data, events of the storage device storing the processed data, etc. The server assigns a unique identifier to each such event, just like attaching a specific label to each event, so that subsequent processing can accurately identify it. Then an initially empty matrix is created, and this matrix is the event stream - knowledge chain association matrix, which will be used to record the association relationships between events, the concerned fault knowledge chain, and each inferential fault knowledge chain.
[0121] Step S122: Traverse each event in the target Internet of Things interaction event stream. For each event, analyze whether the event appears in the concerned fault knowledge chain or any inferred fault knowledge chain. If so, mark the corresponding position in the event stream - knowledge chain association matrix as 1; otherwise, mark it as 0.
[0122] For each event, analyze whether the event appears in the concerned fault knowledge chain or any inferred fault knowledge chain. In a radio monitoring facility, taking the event of sensor data acquisition as an example, if the event of sensor data acquisition is included in the concerned fault knowledge chain (assuming the concerned fault knowledge chain is related to data accuracy) or a certain inferred fault knowledge chain (such as an inferred fault knowledge chain about the impact of sensor hardware failure on data acquisition), then mark the corresponding position in the event stream - knowledge chain association matrix as 1, indicating an association; if it is not in any relevant knowledge chain, mark it as 0. For example, for the data verification event in the data processing unit, if it is related to the currently concerned fault knowledge chain or a certain inferred fault knowledge chain, mark it as 1 at the corresponding position in the matrix; if it is not relevant, mark it as 0.
[0123] Step S123: Use the time window division or event type clustering method to divide the target Internet of Things interaction event stream into multiple stages. For each stage, extract the key features of the events in this stage, and encode the key features into numerical data. After forming the multi - stage encoding features of this stage, combine the multi - stage encoding features of all stages in sequence to obtain the overall multi - stage encoding features of the target Internet of Things interaction event stream. The key features include event type, event frequency, and time interval between events.
[0124] For example, in terms of time window division, during the operation of a radio monitoring facility, the entire interaction event stream can be divided at fixed time intervals, such as every hour as a time window. For each stage, extract the key features of the events in this stage, and encode the key features into numerical data. After forming the multi - stage encoding features of this stage, then combine the multi - stage encoding features of all stages in sequence to obtain the overall multi - stage encoding features of the target Internet of Things interaction event stream. For example, within a time window, the key features of the event of sensor data acquisition include event type (acquiring radio signals), event frequency (such as acquiring once every 10 minutes), and time interval between events (10 - minute interval between each acquisition), etc. Convert these features into numerical data according to the pre - set encoding rules. For example, the event type can be represented by a specific number, the event frequency can be encoded according to the actual frequency value or the difference from the standard frequency, and the time interval between events is also encoded in a similar quantitative way. Perform such operations for each stage, and then combine the encoding features of each stage in sequence to obtain the overall multi - stage encoding features.
[0125] Step S124: For each inferred fault knowledge chain, determine the knowledge chain nodes associated with the events in the target Internet of Things interaction event stream according to the event stream - knowledge chain association matrix.
[0126] Suppose there is an inferred fault knowledge chain regarding the impact of sensor failure on data transmission. According to the event stream - knowledge chain association matrix, if events such as sensor startup event and sensor data sending event are associated with this inferred fault knowledge chain, then the nodes corresponding to these events are the associated knowledge chain nodes in this inferred fault knowledge chain.
[0127] Step S125: For each associated knowledge chain node, perform knowledge chain coding mapping according to the position and role of the knowledge chain node in the knowledge chain and the multi - stage coding features of the corresponding event. The mapping process uses the method of neural network mapping to transform the multi - stage coding features of the corresponding event into mapping coding features corresponding to the knowledge chain node. Also, for each inferred fault knowledge chain, combine the mapping coding features of all associated knowledge chain nodes in sequence to form the multi - stage mapping coding features corresponding to this inferred fault knowledge chain.
[0128] For example, in the inferred fault knowledge chain of sensor failure affecting data transmission, the sensor startup node may be in the starting position in the knowledge chain and has a certain impact on subsequent fault propagation. If the multi - stage coding features of the sensor startup event include startup time, self - check parameters at startup, etc., the neural network mapping will, according to the meaning of the sensor startup node in the knowledge chain, transform these multi - stage coding features into mapping coding features suitable for representing this node in the knowledge chain. For each inferred fault knowledge chain, combine the mapping coding features of all associated knowledge chain nodes in sequence to form the multi - stage mapping coding features corresponding to this inferred fault knowledge chain.
[0129] Step S126: Input each multi - stage mapping coding feature into a recurrent neural network so that the recurrent neural network can generate a fault logic diagram coding vector by learning and extracting the association relationships and temporal information between each multi - stage mapping coding feature.
[0130] For an inferred fault knowledge chain related to an interaction event stream involving multiple links such as sensors, data processing units, and storage devices, the recurrent neural network will analyze the relationships between the multi - stage mapping coding features of events related to different stages and different devices. For example, there may be a temporal sequence relationship and a logical causal relationship between the multi - stage mapping coding features of sensor data acquisition and the multi - stage mapping coding features of data processing by the data processing unit. The recurrent neural network can learn these relationships and integrate them into the fault logic diagram coding vector, which can comprehensively reflect the fault logic relationship in the entire inferred fault knowledge chain.
[0131] Step S127: Analyze the fault propagation path and fault penetration nodes of each inference fault knowledge chain. The fault penetration node refers to a key node that can affect or trigger other faults when a fault propagates in the inference fault knowledge chain.
[0132] Taking the fault propagation in the data processing unit as an example, if an algorithm error in the data processing unit is the fault source, this error may affect subsequent data verification, data storage, etc. along the data processing flow, which is the fault propagation path. The fault penetration node refers to a key node that can affect or trigger other faults when a fault propagates in the inference fault knowledge chain. For example, the data verification module in the data processing unit is a fault penetration node because once it fails, it may cause incorrect data to flow into the storage device, thus triggering more faults.
[0133] Step S128: According to the position, role, and association relationship with other fault penetration nodes of the fault penetration node in the inference fault knowledge chain and the concerned fault knowledge chain, extract the key features of the fault penetration node, and encode the key features as numerical data to form a fault penetration node vector. The key features include node type, node degree, and path length between nodes.
[0134] Taking the data verification module in the data processing unit, which is a fault penetration node, as an example, the node type is the data verification module, which is a specific functional type; the node degree represents the number of edges connected to this node, reflecting its connection relationship with other nodes. For example, if it is connected to both the data processing module and the storage module, then the node degree is 2; the path length between nodes, such as the shortest path length from it to the storage module, etc. Converting these key features into numerical data according to specific coding rules forms the fault penetration node vector.
[0135] Step S129: For each inference fault knowledge chain, stack the corresponding fault logic diagram encoding vector and the fault penetration node vector to obtain the fault logic flow vector corresponding to this inference fault knowledge chain. The attention mechanism or gating mechanism is used in the stacking process to dynamically adjust the weights of the fault logic diagram encoding vector and the fault penetration node vector.
[0136] Finally, taking the inference fault knowledge chain related to the sensor as an example, the fault logic diagram coding vector contains the fault logic relationship in the process of the sensor collecting and transmitting data, and the fault penetration node vector contains the key features of the key components inside the sensor (such as the sensor element, which is a fault penetration node). When stacking, the attention mechanism or the gating mechanism will dynamically allocate weights according to the actual situation. If currently focusing on the impact of the fault of the sensor element on the overall fault logic flow, the weight of the fault penetration node vector may be relatively high; if more attention is paid to the overall fault logic relationship, the weight of the fault logic diagram coding vector may dominate. Through such stacking, the obtained fault logic flow vector can more comprehensively reflect the fault flow situation in the inference fault knowledge chain, providing richer information for subsequent fault analysis and processing.
[0137] In a possible implementation manner, the target fault root cause description label is generated by a first deep learning network, and the training steps of the first deep learning network include:
[0138] Step S101: Obtain the template Internet of Things interaction event stream and the second fault knowledge chain corresponding to the template fault activity in the template Internet of Things interaction event stream, perform fault knowledge chain extraction on the template Internet of Things interaction event stream, and generate the first fault knowledge chain corresponding to each fault interaction activity in the template Internet of Things interaction event stream. The template fault activity is one of the multiple fault interaction activities.
[0139] Step S102: Extract the second interaction path vector of the template Internet of Things interaction event stream, perform encoding representations on the template Internet of Things interaction event stream respectively according to the second fault knowledge chain and each of the first fault knowledge chains, generate multiple template coding features, and extract the template fault penetration node vectors corresponding to the second fault knowledge chain and each of the first fault knowledge chains.
[0140] Step S103: Stack the second interaction path vector, the template coding features, and the template fault penetration node vector and load them into the first deep learning network for prediction to generate a fault root cause description prediction result, which is used to determine the fault root cause description label of the template fault activity.
[0141] Step S104: Obtain the first fault root cause description label of the template fault knowledge point associated with the template Internet of Things interaction event stream, determine the training error according to the fault root cause description prediction result and the first fault root cause description label, and optimize the first deep learning network according to the training error. The template fault knowledge point is used to guide the positioning of the template fault activity.
[0142] In a possible implementation, the template Internet of Things interaction event stream, the second fault knowledge chain, and the first fault root cause description label are all obtained from a sample database. Before step S101, the method further includes:
[0143] Step A110: Obtain a plurality of basic Internet of Things interaction event streams and the fault attention knowledge data corresponding to each of the basic Internet of Things interaction event streams. Based on each of the basic Internet of Things interaction event streams and the corresponding fault attention knowledge data, respectively determine the key interaction point data corresponding to each of the basic Internet of Things interaction event streams. The fault attention knowledge data is used to guide the positioning of key fault activities in the corresponding basic Internet of Things interaction event stream.
[0144] Step A120: Obtain a plurality of reference fault knowledge points. According to each of the key interaction point data, respectively determine the associated fault knowledge points corresponding to each of the basic Internet of Things interaction event streams among the plurality of reference fault knowledge points.
[0145] Step A130: According to each of the basic Internet of Things interaction event streams and the corresponding fault attention knowledge data, respectively determine the fault knowledge chain label data corresponding to each of the basic Internet of Things interaction event streams. The fault knowledge chain label data is used to guide the positioning of the corresponding key fault activities.
[0146] Step A140: Integrate and add each of the basic Internet of Things interaction event streams, the corresponding fault knowledge chain label data, and the corresponding associated fault knowledge points to the sample database. The template Internet of Things interaction event stream is randomly selected from the plurality of basic Internet of Things interaction event streams. The second fault knowledge chain is the fault knowledge chain label data corresponding to the template Internet of Things interaction event stream. The template fault knowledge point is the associated fault knowledge point associated with the template Internet of Things interaction event stream.
[0147] In a possible implementation, step A130 includes:
[0148] Step A131: Load each of the fault attention knowledge data into a second deep learning network for prediction to generate the fault summary description data corresponding to each of the basic Internet of Things interaction event streams.
[0149] Step A132: Respectively perform fault activity positioning according to each of the basic Internet of Things interaction event streams and the corresponding fault summary description data to generate the basic fault logic path corresponding to each of the basic Internet of Things interaction event streams. The basic fault logic path is used to guide the positioning of the corresponding key fault activities.
[0150] Step A133: Load each of the basic IoT interaction event streams and the corresponding basic fault logic paths into the first fault knowledge chain extraction network for fault knowledge chain extraction, generating fault knowledge chain tag data corresponding to each of the basic IoT interaction event streams.
[0151] In this embodiment, first, in the training step of the first deep learning network, the server needs to obtain the template IoT interaction event stream and the second fault knowledge chain corresponding to the template fault activity in the template IoT interaction event stream, and perform fault knowledge chain extraction on the template IoT interaction event stream to generate the first fault knowledge chain corresponding to each fault interaction activity in the template IoT interaction event stream. In a radio monitoring facility, the template IoT interaction event stream can be regarded as a representative interaction event set selected from the IoT interaction event streams in numerous actual operation and maintenance scenarios. For example, this template IoT interaction event stream may include a series of interaction events between sensor devices and data processing units and storage devices within a specific time period. The template fault activity, such as the data collected by the sensor showing abnormal fluctuations at a certain moment, the corresponding second fault knowledge chain of this template fault activity may include a series of fault factors and their associations from the hardware level of the sensor (such as component aging, external interference, etc.) to the software level (such as data collection algorithm errors, etc.), and then to the interaction level with other devices (such as problems with the transmission protocol between the data processing unit and the sensor).
[0152] The server performs fault knowledge chain extraction on this template IoT interaction event stream to generate the first fault knowledge chain corresponding to each fault interaction activity. For example, in addition to the above template fault activity of abnormal sensor data fluctuations, there may also be a fault interaction activity of data processing unit processing data with a delay. For this fault interaction activity of the data processing unit processing data with a delay, the server constructs the corresponding first fault knowledge chain by analyzing each event in the event stream, such as the time when the data processing unit receives data, the duration of processing data, the execution situation of the processing algorithm, etc. This knowledge chain may involve fault factors such as insufficient hardware performance of the data processing unit, high software algorithm complexity, and resource competition when interacting with other devices, and their mutual relationships.
[0153] Next, the server extracts the second interaction path vector of the template IoT interaction event stream. In the scenario of radio monitoring facilities, this second interaction path vector reflects the interaction order and path relationship among various devices in the template IoT interaction event stream. For example, after the sensor collects data and transmits it to the data processing unit, the data processing unit stores the data in the storage device after processing, and the transmission order and connection relationship among these devices constitute the second interaction path vector. At the same time, the server performs encoded representations of the template IoT interaction event stream respectively according to the second fault knowledge chain and each first fault knowledge chain, generates multiple template coding features, and extracts the template fault penetration node vectors corresponding to the second fault knowledge chain and each first fault knowledge chain.
[0154] For the encoded representation, taking the sensor data collection event as an example, according to the second fault knowledge chain or a certain first fault knowledge chain (such as the fault knowledge chain related to the sensor hardware), the server extracts the key features of the sensor data collection event, such as the collection time, collection frequency, data intensity collected, etc., and encodes these features into numerical data according to specific encoding rules. After encoding multiple such events, they are combined to form the template coding features. For the extraction of the template fault penetration node vector, for example, in the sensor hardware fault knowledge chain, the sensor element may be a fault penetration node. The server extracts its key features according to the position of this node in the knowledge chain (such as being at the starting position, with a greater possibility of being the source of the fault), its role (such as directly affecting the accuracy of the collected data), and its association relationship with other fault penetration nodes (such as having an association relationship with the signal conversion module of the sensor, etc.), such as the node type (sensor element), node degree (the number of connections with other related nodes, assuming it is connected to the signal conversion module and the power supply module, the node degree is 2), and the path length between nodes (the number of data processing links passed by the shortest path to the storage device, etc.). After encoding these key features into numerical data, the template fault penetration node vector is formed.
[0155] Then, the server stacks the second interaction path vector, the template encoded features, and the template fault penetration node vector and loads them into the first deep learning network for prediction, generating a prediction result of the fault root cause description. In the intelligent operation and maintenance scenario of radio monitoring facilities, the first deep learning network is a network model designed to analyze the fault root cause. This stacked vector contains the device interaction path in the template IoT interaction event stream, the encoded features of each fault interaction activity, and the information of the fault penetration nodes. After inputting it into the first deep learning network, the first deep learning network performs analysis and prediction according to its internal algorithms and model structure. For example, if the abnormal fluctuation of sensor data is a template fault activity, the network may predict that it is due to the aging of the sensor components resulting in inaccurate data collection, and certain settings of the transmission protocol during data transmission amplify this inaccuracy, ultimately leading to abnormal data fluctuations. This prediction result is the prediction result of the fault root cause description, which is used to determine the fault root cause description label of the template fault activity.
[0156] After that, the server obtains the first fault root cause description label of the template fault knowledge points associated with the template IoT interaction event stream, and determines the training error based on the prediction result of the fault root cause description and the first fault root cause description label, and optimizes the first deep learning network according to the training error. In the radio monitoring facility scenario, the template fault knowledge points are a set of knowledge closely related to the template fault activities, which are used to guide the location of the template fault activities. For example, the template fault knowledge points clearly indicate that the abnormal fluctuation of sensor data may be related to the combined effect of sensor hardware aging and software algorithms, and this is the first fault root cause description label. The server compares the prediction result with this label. If the prediction result only mentions the aging of the sensor hardware and does not involve the influence of the software algorithm, there is a certain deviation, and this deviation is the training error. According to this training error, the server adjusts the parameters of the first deep learning network, such as the weights and biases of the neural network, to improve the prediction accuracy of the network.
[0157] Before obtaining the template IoT interaction event stream and the second fault knowledge chain corresponding to the template fault activities in the template IoT interaction event stream, the server also needs to perform a series of preparatory work to build a sample database. First, the server obtains multiple basic IoT interaction event streams and the corresponding fault attention knowledge data for each basic IoT interaction event stream. In radio monitoring facilities, the basic IoT interaction event stream is a collection of more extensive IoT device interaction events. For example, the interaction event streams between sensors, data processing units, and storage devices in different regions. The fault attention knowledge data is used to guide the location of the key fault activities in the corresponding basic IoT interaction event stream. For example, the information that the sensor data transmission in a certain area is unstable is the fault attention knowledge data.
[0158] Based on each basic Internet of Things (IoT) interaction event stream and the corresponding fault - concerned knowledge data, the server respectively determines the key interaction point data corresponding to each basic IoT interaction event stream. For example, in the basic IoT interaction event stream where sensor data transmission is unstable, the key interaction point data may be the relevant data at the transmission interface between the sensor and the data - processing unit, such as the fluctuation of the transmission rate, the packet - loss situation of the transmitted data, etc. Then, the server obtains multiple reference fault knowledge points, and based on each key interaction point data, respectively determines the associated fault knowledge points corresponding to each basic IoT interaction event stream in the multiple reference fault knowledge points. A reference fault knowledge point is a set containing various fault types and their related knowledge. According to the key interaction point data, such as transmission rate fluctuation and packet - loss situation, the associated fault knowledge points that match in the reference fault knowledge points may be knowledge about transmission protocol errors or network congestion, etc.
[0159] Then, the server respectively determines the fault knowledge chain label data corresponding to each basic IoT interaction event stream based on each basic IoT interaction event stream and the corresponding fault - concerned knowledge data. Specifically, the server loads each fault - concerned knowledge data into the second deep - learning network for prediction, generating the fault summary description data corresponding to each basic IoT interaction event stream. In the scenario of radio monitoring facilities, taking the unstable sensor data transmission as an example, after the fault - concerned knowledge data (unstable transmission information) is loaded into the second deep - learning network, the network may predict that the data transmission instability is caused by problems in the network transmission link, and this result is the fault summary description data.
[0160] Next, the server respectively conducts fault activity location based on each basic IoT interaction event stream and the corresponding fault summary description data, generating the basic fault logic path corresponding to each basic IoT interaction event stream. For example, according to the fault summary description data of unstable sensor data transmission and problems in the network transmission link, the basic fault logic path may start from the sensor sending data, go through the processing of the transmission protocol, and due to network congestion or transmission protocol errors during network transmission, the data becomes unstable when it reaches the data - processing unit.
[0161] Finally, the server loads each basic Internet of Things interaction event stream and the corresponding basic fault logic path into the first fault knowledge chain extraction network for fault knowledge chain extraction, and generates fault knowledge chain label data corresponding to each basic Internet of Things interaction event stream. This fault knowledge chain label data can guide the positioning of the corresponding key fault activities. For example, in the case of unstable sensor data transmission, the fault knowledge chain label data includes various factors and their mutual relationships from aspects such as sensor hardware, transmission protocol, and network environment. These relationships can clearly indicate which factors are the key factors leading to the key fault activity of unstable data transmission, thus providing important data support for the subsequent construction of the sample database and the training of the first deep learning network.
[0162] Thereby, the first deep learning network can be effectively trained, so as to accurately analyze the root cause of the fault and generate an accurate root cause of the fault description label.
[0163] Figure 2 FIG. shows the hardware structure diagram of the distributed operation and maintenance intelligent management system 100 based on cloud computing and the Internet of Things provided by the embodiment of the present invention for implementing the above-mentioned distributed operation and maintenance intelligent management method based on cloud computing and the Internet of Things, as Figure 2 shown, the distributed operation and maintenance intelligent management system 100 based on cloud computing and the Internet of Things may include a processor 110, a machine-readable storage medium 120, a bus 130, and a communication unit 140.
[0164] The machine-readable storage medium 120 can store data and / or instructions. In some embodiments, the machine-readable storage medium 120 can store data obtained from an external terminal. In some embodiments, the machine-readable storage medium 120 can store data and / or instructions that the distributed operation and maintenance intelligent management system 100 based on cloud computing and the Internet of Things uses to execute or use to complete the exemplary methods described in the present invention.
[0165] In a specific implementation process, one or more processors 110 execute computer-executable instructions stored in the machine-readable storage medium 120, so that the processors 110 can execute the distributed operation and maintenance intelligent management method based on cloud computing and the Internet of Things in the above method embodiments. The processors 110, the machine-readable storage medium 120, and the communication unit 140 are connected through the bus 130, and the processors 110 can be used to control the transceiver actions of the communication unit 140.
[0166] The specific implementation process of the processor 110 can refer to the various method embodiments executed by the above-mentioned distributed operation and maintenance intelligent management system 100. The implementation principles and technical effects are similar, and will not be elaborated here in this embodiment.
[0167] In addition, an embodiment of the present invention further provides a readable storage medium, in which computer-executable instructions are preset. When a processor executes the computer-executable instructions, the distributed operation and maintenance intelligent management method based on cloud computing and the Internet of Things as described above is implemented.
[0168] It should be noted that, in order to simplify the description of the present invention disclosure and thus help the understanding of one or more embodiments of the invention, in the foregoing description of the embodiments of the present invention, sometimes multiple features are merged into one embodiment, drawing, or description thereof.
Claims
1. A distributed operation and maintenance intelligent management method based on cloud computing and the Internet of Things, characterized in that, The method includes: Obtaining inference fault knowledge chains corresponding to respective potential fault activities in a target Internet of Things (IoT) interaction event stream in a cloud computing environment, and determining a concerned fault knowledge chain among the multiple inference fault knowledge chains, where the concerned fault knowledge chain is used to guide the positioning of a target fault activity that is concerned among the multiple potential fault activities; Performing encoding representations of the target IoT interaction event stream respectively according to the concerned fault knowledge chain and each of the inference fault knowledge chains to generate multiple fault logic graph encoding vectors, and extracting fault penetration node vectors corresponding to the concerned fault knowledge chain and each of the inference fault knowledge chains; wherein, a fault penetration node refers to a key node that can affect or trigger other faults when a fault propagates in the inference fault knowledge chain; Stacking each of the fault logic graph encoding vectors with the corresponding fault penetration node vector to generate multiple fault logic transfer vectors; Extracting a first interaction path vector of the target IoT interaction event stream, and stacking the first interaction path vector and the multiple fault logic transfer vectors to generate a target stacked vector; Performing prediction according to the target stacked vector to generate a target fault root cause description label of the target fault activity; The performing encoding representations of the target IoT interaction event stream respectively according to the concerned fault knowledge chain and each of the inference fault knowledge chains to generate multiple fault logic graph encoding vectors includes: Performing multi-stage encoding representation on the target IoT interaction event stream to generate multi-stage encoding features of the target IoT interaction event stream; Performing knowledge chain encoding mapping on the multi-stage encoding features respectively according to the concerned fault knowledge chain and each of the inference fault knowledge chains to generate multiple multi-stage mapping encoding features; Inputting each multi-stage mapping encoding feature into a recurrent neural network so that the recurrent neural network generates a fault logic graph encoding vector by learning and extracting the association relationship and temporal information between each multi-stage mapping encoding feature; The obtaining inference fault knowledge chains corresponding to respective potential fault activities in a target IoT interaction event stream in a cloud computing environment, and determining a concerned fault knowledge chain among the multiple inference fault knowledge chains includes: Obtaining a target IoT interaction event stream and guiding identification information of the target IoT interaction event stream, where the target IoT interaction event stream includes multiple potential fault activities, and the guiding identification information is used to guide the positioning of a target fault activity that is concerned among the multiple potential fault activities; Performing fault knowledge chain extraction on the target IoT interaction event stream to generate inference fault knowledge chains corresponding to the respective potential fault activities, and determining a concerned fault knowledge chain among the inference fault knowledge chains according to the guiding identification information; The step of performing fault knowledge chain extraction on the target IoT interaction event stream to generate inference fault knowledge chains corresponding to the respective potential fault activities includes: Traverse the target Internet of Things interaction event stream, add each event in the target Internet of Things interaction event stream as a node to the event association graph, determine the association relationship between events according to the chronological order, event source and event target attributes between events, and after adding corresponding edges in the event association graph, optimize the event association graph; Use a graph traversal algorithm to traverse the nodes in the optimized event association graph, identify potential fault activities according to the association relationship between events and preset fault identification rules, represent the identified fault activities in the form of a node set, and add them to the list of potential fault activities; For each potential fault activity in the list of potential fault activities, extract the corresponding key features, and the key features include event type distribution, event time interval distribution, event source and target distribution; Quantify the key features into numerical data, and after combining them into a fault feature vector, perform normalization processing on the fault feature vector; For the fault feature vector of each potential fault activity, use a similarity calculation algorithm to find the most similar fault feature in a pre-constructed fault knowledge base, obtain the corresponding fault knowledge data according to the found fault feature, and construct an inference fault knowledge chain according to the association relationship between fault types and the logical relationship between fault causes in the fault knowledge data.
2. The distributed operation and maintenance intelligent management method based on cloud computing and the Internet of Things according to claim 1, wherein The stacking of the first interaction path vector and multiple fault logic transfer vectors to generate a target stacking vector includes: Respectively determine the knowledge chain lengths of each of the inference fault knowledge chains, and stack the fault logic transfer vectors corresponding to each of the inference fault knowledge chains in descending order of the knowledge chain lengths to generate a first stacking vector; Stack the first interaction path vector, the first stacking vector, and the fault logic transfer vector corresponding to the concerned fault knowledge chain to generate a target stacking vector.
3. The distributed operation and maintenance intelligent management method based on cloud computing and the Internet of Things according to claim 2, characterized in that, The target fault root cause description label is generated by a first deep learning network. The stacking of the first interaction path vector, the first stacking vector, and the fault logic transfer vector corresponding to the concerned fault knowledge chain to generate a target stacking vector includes: Generate guiding information for guiding the first deep learning network to generate a fault root cause description label; Extract the guiding knowledge vector of the guiding information, and stack the first interaction path vector, the guiding knowledge vector, the first stacking vector, and the fault logic transfer vector corresponding to the concerned fault knowledge chain to generate a target stacking vector.
4. The distributed operation and maintenance intelligent management method based on cloud computing and the Internet of Things according to claim 1, wherein The target fault root cause description label is generated by a first deep learning network. The training steps of the first deep learning network include: Obtain a template Internet of Things interaction event stream and a second fault knowledge chain corresponding to a template fault activity in the template Internet of Things interaction event stream, extract a fault knowledge chain from the template Internet of Things interaction event stream to generate a first fault knowledge chain corresponding to each fault interaction activity in the template Internet of Things interaction event stream, and the template fault activity is one of the multiple fault interaction activities; Extract the second interaction path vector of the template Internet of Things interaction event stream, perform encoded representations of the template Internet of Things interaction event stream respectively according to the second fault knowledge chain and each of the first fault knowledge chains, generate multiple template encoding features, and extract the template fault penetration node vectors corresponding to the second fault knowledge chain and each of the first fault knowledge chains; Stack the second interaction path vector, the template encoding features, and the template fault penetration node vectors and load them into the first deep learning network for prediction to generate a predicted result of the root cause description of the fault, and the predicted result of the root cause description of the fault is used to determine the root cause description label of the template fault activity; Obtain the first root cause description label of the template fault knowledge point associated with the template Internet of Things interaction event stream, determine the training error according to the predicted result of the root cause description of the fault and the first root cause description label, and optimize the first deep learning network according to the training error. The template fault knowledge point is used to guide the positioning of the template fault activity.
5. The distributed operation and maintenance intelligent management method based on cloud computing and the Internet of Things according to claim 4, characterized in that The template Internet of Things interaction event stream, the second fault knowledge chain, and the first root cause description label are all obtained from the sample database. Before obtaining the template Internet of Things interaction event stream and the second fault knowledge chain corresponding to the template fault activity in the template Internet of Things interaction event stream, the method further includes: Obtain multiple basic Internet of Things interaction event streams and the fault attention knowledge data corresponding to each of the basic Internet of Things interaction event streams. Based on each of the basic Internet of Things interaction event streams and the corresponding fault attention knowledge data, respectively determine the key interaction point data corresponding to each of the basic Internet of Things interaction event streams. The fault attention knowledge data is used to guide the positioning of the key fault activity in the corresponding basic Internet of Things interaction event stream; Obtain multiple reference fault knowledge points, and respectively determine the associated fault knowledge points corresponding to each of the basic Internet of Things interaction event streams among the multiple reference fault knowledge points according to the key interaction point data of each; Based on each of the basic Internet of Things interaction event streams and the corresponding fault attention knowledge data, respectively determine the fault knowledge chain label data corresponding to each of the basic Internet of Things interaction event streams. The fault knowledge chain label data is used to guide the positioning of the corresponding key fault activity; Integrate and add each of the basic Internet of Things interaction event streams, the corresponding fault knowledge chain label data, and the corresponding associated fault knowledge points to the sample database. The template Internet of Things interaction event stream is randomly selected from the multiple basic Internet of Things interaction event streams. The second fault knowledge chain is the fault knowledge chain label data corresponding to the template Internet of Things interaction event stream, and the template fault knowledge point is the associated fault knowledge point associated with the template Internet of Things interaction event stream.
6. The distributed operation and maintenance intelligent management method based on cloud computing and the Internet of Things according to claim 5, characterized in that, The determining, based on each of the basic Internet of Things interaction event streams and the corresponding fault attention knowledge data, of the fault knowledge chain label data corresponding to each of the basic Internet of Things interaction event streams includes: Load each of the fault - concerned knowledge data into the second deep - learning network for prediction, and generate fault summary description data corresponding to each of the basic Internet - of - Things interaction event streams; Perform fault activity localization respectively based on each of the basic Internet - of - Things interaction event streams and the corresponding fault summary description data, and generate basic fault logic paths corresponding to each of the basic Internet - of - Things interaction event streams, where the basic fault logic paths are used to guide the localization of the corresponding key fault activities; Load each of the basic Internet - of - Things interaction event streams and the corresponding basic fault logic paths into the first fault knowledge chain extraction network for fault knowledge chain extraction, and generate fault knowledge chain label data corresponding to each of the basic Internet - of - Things interaction event streams.
7. A distributed operation and maintenance intelligent management system based on cloud computing and the Internet of Things, characterized in that, The distributed operation and maintenance intelligent management system based on cloud computing and the Internet of Things includes a processor and a memory. The memory is connected to the processor. The memory is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the memory to implement the distributed operation and maintenance intelligent management method according to any one of claims 1 - 6 above.
Citation Information
Patent Citations
Processing method and device for abnormal data of service system, equipment and storage medium
CN113687972A
Fault root cause positioning method and device based on knowledge graph, equipment and medium
CN114218403A